Qwen3.6 MTP and API / Connections #5556
danielhanchen
announced in
Announcements
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
2x faster GGUF via auto enabled MTP, API calling to OpenAI, Anthropic, OpenRouter with auto prompt caching, web search, code execution, connecting to external vLLM, Ollama, llama-server, security, UI, UX updates, Experimental MLX inference and more!
MTP speculative decoding support 1.4 to 2x faster inference!
API provider calling & external connections
MLX inference (Experimental)
Other Unsloth Studio updates
dir="auto", long log-line truncation fixTraining updates
grouped_mmcontraction crashlm_head+ cross-entropy forward, with single-matmul path underUNSLOTH_RETURN_LOGITS=1HF_DATASETS_OFFLINEalongsideHF_HUB_OFFLINEUnsloth Studio security improvements
hf upload,NOFILE)torch.loadfallback ontraining_args.binso untrusted pickles can never execute on model loadBug fixes and correctness
num_logits_to_keepregression fixed on transformers >= 4.52fast_generateunifies legacy and new logits kwargs (fixes Mistral merge site)higher_precision_softmaxmade idempotentLOSS_MAPPINGkey aliased toForCausalLMLoss(covers transformers 5.x)unsloth-runextra args/recommended-folders500 on unreadable model directories under Python 3.12+Installer and platform reliability
STUDIO_HOME/UNSLOTH_STUDIO_HOMEggml-org/llama.cppprebuiltscudartbundle and Torch NVIDIA DLL paths added toPATHflash-attninstall on Blackwell GPUs (sm_100+)--gcc-install-diruninstall.sh) and Windows (uninstall.ps1)unsloth --versionflagWhat's Changed in Unsloth
with:keys in notebooks-ci checkout steps by @rolandtannous in ci: merge duplicatewith:keys in notebooks-ci checkout steps #5447cache: 'npm'from setup-node (silent abort on Windows) by @danielhanchen in ci: dropcache: 'npm'from setup-node (silent abort on Windows) #5474save_pretrained_merged#5410) by @danielhanchen in tests: pinned-symbol canary for unsloth-zoo save_pretrained_merged guards (#5410) #5433/v1/chat/completionsdefaults to streaming whenstreamis omitted #5047) by @wtfashwin in studio/openai: align chat completions docstring with stream=false default (closes #5047) #5524New Contributors
Full Changelog: v0.1.39-beta...v0.1.40-beta
What's Changed in Unsloth-Zoo
chat_templateusing tuple raise reference before assignment on Ollama #637 landing by @danielhanchen in chore: trim verbose comments across PR #637 landing unsloth-zoo#640save_pretrained_merged#5410) by @danielhanchen in saving: layout-aware MoE LoRA merge + loud-fail on fallback (#5410) unsloth-zoo#647save_pretrained_merged#5410) by @danielhanchen in tests: CPU regression detectors for the MoE merge / save path (#5410) unsloth-zoo#655This discussion was created from the release Qwen3.6 MTP and API / Connections.
All reactions