Ad
 
Learn More
Favicon of Transformers Releases

Transformers Releases

Active

Hugging Face • Last updated about 2 hours ago

Activity Score

90/100
Recency:30/30
Cadence:20/30
Completeness:30/30
Health:10/10
  • Updated in the last week
  • 4 updates in last 30 days
  • Complete entries with dates, titles, URLs, and summaries

Recent Updates

Release: v5.16.0

Release v5.16.0 New Model additions Qwen4-Exp Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE). GR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm. It mixes multiple residual streams with fine-grained elementwise gating before each attention and Mixture-of-Experts (MoE) block, then controls how much of the block output is injected back into each stream. QSA uses multiple query heads to score compressed key blocks, selects the most relevant contiguous token blocks, and keeps the incomplete trailing block uncompressed. This block-level selection reduces indexing overhead and improves memory locality for long sequences. Combined with Gated DeltaNet, QSA makes Qwen4-Exp the first hybrid architecture to integrate linear and sparse attention, substantially improving inference efficiency for long-context workloads. PLE enriches selected decoder layers with layer-specific lexical features derived from hashed token n-grams and a dilated depthwise convolution. Links: Documentation Add Qwen4Exp model ( #48337 ) by @Cyrilvallez in #48337 GraniteSpeech5 Granite Speech 5.0 Turbo CTC is a lightweight (~470M parameters) conformer encoder for automatic speech recognition, trained with Connectionist Temporal Classification (CTC) on BPE targets. It is a fast, encoder-only member of the Granite Speech family: transcription requires a single forward pass followed by greedy CTC decoding, with no autoregressive decoder. Architecturally, it extends the Granite Speech conformer CTC encoder with: Frame stacking + block-wise time subsampling : the feature extractor stacks pairs of log-mel(+delta) frames (2x), and the first two conformer blocks each subsample time by 2 through a stride-2 depthwise convolution (with a mean-pooled residual), for a total 8x time reduction at 10 ms mel hop. Block attention with Shaw's relative positional embeddings : attention is computed over fixed-size blocks (the sequence is right-padded to a whole number of blocks, with padded frames masked out), using separate bias-free query/key/value projections. Self-conditioned CTC : the CTC posteriors of the middle layer are projected and fed back into the hidden states, and the CTC head is shared between this mid-layer self-conditioning and the final prediction. Links: Documentation Add Granite Speech 5.0 - ( #48288 ) by @eustlb in #48288 Step3p7 Step-3.7-Flash was proposed in Step 3.7 Flash by StepFun. It is a 198B-parameter sparse Mixture-of-Experts vision-language model, pairing a 196B-parameter MoE language backbone with a 1.8B-parameter vision encoder for native image understanding. StepFun hasn't published a technical report for Step-3.7-Flash, so the details below are drawn from the released checkpoint's configuration rather than a paper. Sparse MoE decoder : all but the first 3 decoder layers route through a MoE block of 288 routed experts (top-8 per token) plus a single shared expert. The router scores experts with a sigmoid and a learned per-expert bias instead of an auxiliary load-balancing loss, the same strategy as DeepSeek-V3 . Gated attention : each attention layer adds an extra projection whose sigmoid output gates the attention output per head, before the output projection — the same Gated Attention mechanism used in Qwen3-Next . A subset of layers use fewer heads and a sliding window instead of full attention. Multi-token prediction : some checkpoints ship extra decoder layers trained for multi-token prediction, which [ ~GenerationMixin.generate ] can use for speculative decoding via use_mtp=True . Vision encoder : a SigLIP-style ViT with 2-D rotary position embeddings and a learned per-layer scale on the attention and MLP branches. Its output is downsampled 4x by two stride-2 convolutions before a linear projector maps it into the text model's hidden size. Dynamic image tiling : instead of a fixed tile grid, the image processor picks its tiling window from each image's own aspect ratio, producing one downscaled global view plus zero or more local high-resolution crops per image. Links: Documentation [new model] step 3.7 ( #46658 ) by @itazap in #46658 CohereCompass CohereCompass is the base architecture for small, specialized (vision-)language models trained by Cohere. Links: Documentation Add CohereCompass modeling ( #47878 ) by @calpt in #47878 ESMC and ESMFold2 ESMC and ESMFold2 are new state-of-the-art protein language and folding models from BioHub. ESMC is trained with a masked language modeling objective, and it can be easily transferred to sequence and token classification tasks for proteins. Checkpoints exist in various sizes, from 300M parameters up to 6B parameters. It works as a drop-in replacement for older ESM-2 and ESM-3 models, with significantly higher accuracy. ESMFold2 is a state-of-the-art protein folding model which produces high accuracy predictions. It uses an iterated diffusion approach that is significantly different from the original ESMFold, offering huge improvements in accuracy for more complex structures. Links: Documentation ESMC , Documentation ESMFold2 Port ESMC and ESMFold2 to Transformers ( #46419 ) by @Rocketknight1 in #46419 Breaking changes The legacy tensor-parallel implementation has been replaced with a DTensor-native backend, so users relying on the previous TP API for inference or training must migrate to the new DTensor-based interface. 🚨 TP dtensor API inference + training ( #47579 ) by @3outeille attn_implementation="sdpa" dispatch is now properly supported for wav2vec2_conformer, wav2vec2-bert, and SeamlessM4T/v2 models, which may change initialization behavior for users who previously worked around this limitation. 🚨[wav2vec2] Support attn_implementation=sdpa dispatch ( #46196 ) by @YangKai0616 FuyuProcessor no longer returns the image_patch_indices output, so any code that depends on this field must be updated to remove references to it. 🚨 Leftover processors ( #47924 ) by @zucchini-nlp Cache Several cache-related bugs were fixed in this release, including an off-by-one error in the sliding window cache, Whisper speculative decoding cache corruption, CpmAnt use-cache failures, Qwen2.5-Omni/Qwen3-Omni-MoE generation with compilable caches, and compressed-tensors loading for KV-cache-only quantized models. Documentation was also added for cache token removal using negative values, and per-layer cache configuration support (allowing models to use different cache settings per layer) was introduced. Cpmant fix use cache ( #48013 ) by @jiqing-feng in [ #48013 ] [docs] Cache crop ( #47950 ) by @stevhliu in [ #47950 ] Revert "Support per-layer cache configuration and attention-mask selection" ( #48175 ) by @Cyrilvallez in [ #48175 ] Support per-layer cache configuration and attention-mask selection ( #47901 ) by @eladsegal in [ #47901 ] Fix Qwen2.5-Omni / Qwen3-Omni-MoE generation with a compilable cache ( #47872 ) by @jiqing-feng in [ #47872 ] Fix sliding window cache index off-by-one on wraparound ( #47708 ) by @hameedibrh in [ #47708 ] [Whisper] Fix speculative decoding: UnboundLocalError, cache corruption, and speed regression ( #48000 ) by @ydshieh in [ #48000 ] Fix compressed-tensors loading for KV-cache-only quantized models ( #47904 ) by @kylesayrs in [ #47904 ] Generation This release fixes several generation bugs across multiple models, including Whisper speculative decoding issues (UnboundLocalError, cache corruption, speed regression, and left-padded batch position IDs), broken image generation in Emu3, garbage output in OLMo/GPTNeoX, and Qwen2.5-Omni/Qwen3-Omni-MoE generation with compilable caches. Additionally, logit distributions for candidate generators using sampling are now aligned by returning logits after applying logit processors. [Whisper] Fix speculative decoding: preserve cleared suppress tokens through super().generate() ( #48108 ) by @ydshieh in [ #48108 ] [Whisper] Fix decoder position IDs for left-padded batches in longform generation ( #48028 ) by @ydshieh in [ #48028 ] [serge] Fix 2 integration tests for model generation failing with import_or_config (other (2)) ( #48061 ) by @sergereview[bot] in [ #48061 ] Align logit distributions for CandidateGenerators using sampling ( #48007 ) by @Cyrilvallez in [ #48007 ] [GPTNeoX] Fix post_processor not overridden when loading from pretrained (OLMo garbage generation) ( #47988 ) by @ydshieh in [ #47988 ] [emu3] 🦮 Black Labrador is back! Fix image generation broken since #37033 ( #47948 ) by @ydshieh in [ #47948 ] Attention Several attention-related bug fixes were made in this release, including correcting a SigLIP2 documentation typo, fixing Flash/SDPA attention dispatch tests for xcodec2 and ROCm RDNA GPUs, resolving a GPT2 cross-attention mask being silently discarded, and enabling SDPA support declaration in TimmWrapper . Per-layer cache configuration and attention-mask selection support was also introduced, allowing models with heterogeneous layer configurations to use distinct sliding_window , attention_chunk_size , and number_of_conv_states values per layer. doc: Fix typo in SigLIP2 Flash Attention code example ( #48197 ) by @VimalN2005 in [ #48197 ] [xcodec2] Fix flex attention and flash dispatch tests ( #48244 ) by @jiqing-feng in [ #48244 ] Fix ROCm SDPA-flash skip guard that crashes on RDNA GPUs ( #47965 ) by @Abdennacer-Badaoui in [ #47965 ] [GPT2] Fix encoder_attention_mask being silently discarded in cross-attention ( #47946 ) by @DavidJohnQuinlan in [ #47946 ] Declare sdpa support in TimmWrapper ( #47939 ) by @jiqing-feng in [ #47939 ] Quantization Quantization improvements include adding NVFP4 quantization support via HF kernels (enabling on-the-fly BF16 weight quantization with ~50% memory reduction), and fixing several bugs: reverting a regression in is_quantization_compressed that caused incorrect module layouts for packed-format checkpoints, fixing CLIP weight initialization failures with quantized checkpoints, and restoring KV-cache quantization setup for KV-cache-only quantized models. Revert "[Quantization]: Refactor is_quantization_compressed for format-based detection" ( #48072 ) by @subin9 in [ #48072 ] feat: add nvfp4 quantization ( #47883 ) by @drbh in [ #47883 ] [DeepSeekV2] Fix integration tests OOM: use device_map=auto instead of 8-bit quantization ( #47991 ) by @ydshieh in [ #47991 ] Fix CLIP _init_weights when a child module carries quantized weights ( #47921 ) by @Bluear7878 in [ #47921 ] Parallelization Introduced a naive pipeline parallel inference engine supporting tied/untied weight embeddings with seamless generate() integration, while restoring backward compatibility for the tensor-parallel API with a deprecation cycle for tp_plan in from_pretrained() . Additionally fixed a model parallel bug in the BLT model affecting beam search. Restore BC for the tensor-parallel API ( #48300 ) by @ArthurZucker in [ #48300 ] fix bug for blt model parallel bug ( #48327 ) by @kaixuanliu in [ #48327 ] Pipeline parallel naive inference ( #47289 ) by @3outeille in [ #47289 ] Kernels Kernel support was improved with documentation updates highlighting supported models, a fix for export crashes on kernel-decorated functions by adding a is_torchdynamo_exporting guard, and the default Flash Attention 2 hub kernel version was bumped to v3 to resolve compatibility issues with newer PyTorch versions. [docs] Kernel supported models ( #48258 ) by @stevhliu in [ #48258 ] [Fix] Export crashes on kernel-decorated function ( #47808 ) by @remi-or in [ #47808 ] Bump default flash-attn2 hub kernel version to v3 ( #47863 ) by @jiqing-feng in [ #47863 ] Bugfixes and improvements Fix video-llama modular conversion ( #48336 ) by @zucchini-nlp in [ #48336 ] CI: gate the hunyuan-moe slow test ( #48330 ) by @tarekziade in [ #48330 ] Add a regression test for force_accelerate_hooks signature preservation ( #48260 ) by @wtdcode in [ #48260 ] Add an opt-in per-frame pixel cap (cap_pixels_per_frame) to the Qwen3-VL video processor ( #48071 ) by @dkrisman in [ #48071 ] Fix scores type in stopping criteria docstrings ( #47676 ) by @qgallouedec in [ #47676 ] Docstring check didn't match some file - fix it ( #48121 ) by @zucchini-nlp in [ #48121 ] Add shared ImageProcessingTester ( #47745 ) by @guarin in [ #47745 ] Fix AutoTokenizer returning TokenizersBackend for DeepSeek-R1-Distill-Qwen models ( #48211 ) by @ydshieh in [ #48211 ] [docs] Fix failing doctests ( #47687 ) by @stevhliu in [ #47687 ] Fix build_2d_sinusoidal_position_embedding on MPS ( #47897 ) by @guarin in [ #47897 ] [Gemma4] Investigate flaky test_generation_beyond_sliding_window_1_eager ( #48236 ) by @ydshieh in [ #48236 ] Let gradient checkpointing skip layers with every_n_layers ( #48200 ) by @qgallouedec in [ #48200 ] Disable daily nightly CI ( #48292 ) by @remi-or in [ #48292 ] gs ( #48288 ) by @eustlb in [ #48288 ] CI: fix muse OOMs ( #48284 ) by @tarekziade in [ #48284 ] Fix BayesianDetectorModel.from_pretrained() by calling post_init() ( #48254 ) by @woojinpaik in [ #48254 ] [Fix] Avoid duplicating tests in CI ( #48287 ) by @remi-or in [ #48287 ] [ GDN ] Fix recurrent FLA fallback ( #48266 ) by @vasqu in [ #48266 ] Fix tie_word_embeddings not lifted from text_config for some VLM configs (BC regression) ( #45857 ) by @qgallouedec in [ #45857 ] replace xpu-smi subprocess call in benchmark_v2 ( #48083 ) by @kaixuanliu in [ #48083 ] ignore mlinter ci file ( #48267 ) by @tarekziade in [ #48267 ] Fix dtype mismatch in grouped_mm_fallback for LoRA training on Mamba+… ( #47933 ) by @adh-aakriti in [ #47933 ] Compute MoE load-balancing loss per layer to avoid giant one-hot materialization −99.7% @ 128k ( #48131 ) by @qgallouedec in [ #48131 ] deterministic layer_types buffer registration in multiple models ( #48162 ) by @mowoe in [ #48162 ] Fix nemotron_h save_pretrained emitting singular backbone.embedding.weight ( #48075 ) by @yuekaizhang in [ #48075 ] force_accelerate_hooks should not hide the signature it wraps ( #48156 ) by @SunMarc in [ #48156 ] fix(data_collator): align TokenClassification numpy_call with torch_call ( #48212 ) by @ in [ #48212 ] Fix a typo in a use of a local variable field_ in a test ( #48184 ) by @AleksMat in [ #48184 ] [serge] Fix 2 integration tests regressed by commit 16780c8 (PR #47622 ) ( #48134 ) by @sergereview[bot] in [ #48134 ] [Gemma4] Fix stale expected values in integration tests ( #48233 ) by @ydshieh in [ #48233 ] [TableTransformer, PI0] Fix stale expected values and OOM in integration tests ( #48198 ) by @ydshieh in [ #48198 ] Post two CI badges on a PR: CPU PR CI and GPU run-slow ( #48190 ) by @tarekziade in [ #48190 ] Fix stale expected values in integration tests (cuda sm_86 / Aug04 regressions) ( #48171 ) by @ydshieh in [ #48171 ] Assign a reviewer even when a codeowner has left, and route models by modality ( #48085 ) by @tarekziade in [ #48085 ] [VITS] Un-skip test_model_forward ( #46375 ) by @blipbyte in [ #46375 ] Apply context parallelism to the evaluation path ( #48167 ) by @qgallouedec in [ #48167 ] Fix gpt_oss runs on GPU ( #48118 ) by @tarekziade in [ #48118 ] [EsmFold2] Fix stale expected distogram logit values ( #48182 ) by @ydshieh in [ #48182 ] Fix Apr 05 integration test regressions (cuda sm_86) ( #48170 ) by @ydshieh in [ #48170 ] Add MLU support to is_flash_linear_attention_available ( #46995 ) by @atri2549 in [ #46995 ] Fix integration test expected values for cuda sm_86 (Mar 15 regressions) ( #48168 ) by @ydshieh in [ #48168 ] Fix DynamicCache reconstruction during ExecuTorch export ( #47900 ) by @eladsegal in [ #47900 ] [LLaVA] Fix pixtral integration tests for cuda sm_86 ( #48166 ) by @ydshieh in [ #48166 ] Port ESMC and ESMFold2 to Transformers ( #46419 ) by @Rocketknight1 in [ #46419 ] [Qwen2.5-Omni] Update stale expected values for cuda sm_86 ( #48164 ) by @ydshieh in [ #48164 ] [Mistral3] Fix batched integration tests: padding_side=left + update expected values ( #48161 ) by @ydshieh in [ #48161 ] Fix DeepSeek V2 default vocab size ( #48159 ) by @hmellor in [ #48159 ] [InternVL] Fix stale expected values for Llama integration tests (cuda sm_80) ( #48153 ) by @ydshieh in [ #48153 ] Retry transient network errors (RemoteDisconnected) in github_utils ( #48124 ) by @ydshieh in [ #48124 ] fix bugs for clvp model ( #47127 ) by @kaixuanliu in [ #47127 ] Always tie embeddings for LongT5 and Pop2Piano ( #47620 ) by @jiqing-feng in [ #47620 ] Use generator with seed for LengthGroupedSampler in Trainer._get_eval_sampler for deterministic eval order with per_device_eval_batch_size > 1 ( #48025 ) by @philipshurpik in [ #48025 ] Enable mlinter findings artifact for inline PR reviews ( #48117 ) by @ydshieh in [ #48117 ] Fix CpmAnt loading: size lm_head to vocab_size ( #48012 ) by @jiqing-feng in [ #48012 ] Delete old mlinter review comments before posting new ones ( #48107 ) by @ydshieh in [ #48107 ] [Video] Warn and return all frames when num_frames exceeds total_num_frames ( #48074 ) by @carlszk in [ #48074 ] Accept artifact dir as argument in post_mlinter_review.py ( #48106 ) by @ydshieh in [ #48106 ] Fix mlinter artifact path ( #48088 ) by @ydshieh in [ #48088 ] [Video] Fix convert_to_rgb channel slicing and alpha blending for RGBA videos ( #48053 ) by @ in [ #48053 ] [PE] Skip test_sdpa_can_dispatch_on_flash for TimmWrapper-backed models ( #48064 ) by @ydshieh in [ #48064 ] fix(pipeline): preserve model.generation_config precedence over pipeline defaults ( #47752 ) ( #47953 ) by @nithin42 in [ #47953 ] [Gemma3] Update integration test expected values for A10G ( #48036 ) by @ydshieh in [ #48036 ] Fallback from 'lanczos' to 'bicubic' when on cuda ( #48026 ) by @zucchini-nlp in [ #48026 ] [serge] Fix 2 integration tests regressed by commit b9090ae (PR #47096 ) ( #48060 ) by @sergereview[bot] in [ #48060 ] [Fix] Small FA-related test failures in CB ( #47341 ) by @remi-or in [ #47341 ] fix: honor empty processor_kwargs={} in multimodal pipelines ( #48044 ) by @ in [ #48044 ] [CircleCI] Enable CI for private forks, no-op for public repo ( #48056 ) by @ydshieh in [ #48056 ] Support BatchFeature in length-grouped samplers ( #48034 ) by @qgallouedec in [ #48034 ] Fix EOS for candidate generators ( #47931 ) by @ in [ #47931 ] [Gemma3n] Update integration test expected values for A10G + torch 2.13 ( #48035 ) by @ydshieh in [ #48035 ] fix failed test cases for muse_glimmer ( #48011 ) by @kaixuanliu in [ #48011 ] Cohere compass tests ( #47895 ) by @zucchini-nlp in [ #47895 ] docs: use relative paths for README language menus and add fa/ro entries ( #47777 ) by @Priyans-Lathiya in [ #47777 ] Moving mlinter to 0.1.4 ( #47918 ) by @tarekziade in [ #47918 ] [MoE] Fix Blackwell GPU crash with torch._grouped_mm on torch <= 2.8 ( #48014 ) by @ in [ #48014 ] [docs] Muse Glimmer ( #47882 ) by @stevhliu in [ #47882 ] [Florence2] Fix two integration test failures caused by torch 2.13 and auto-dtype ( #48031 ) by @ydshieh in [ #48031 ] Proper separation of tests ( #47943 ) by @zucchini-nlp in [ #47943 ] unpin pytest in the examples_torch deps ( #48023 ) by @tarekziade in [ #48023 ] [CI] Fix startup failure in pr_build_doc_with_comment workflow by adding missing get-pr-number dependency ( #47971 ) by @ in [ #47971 ] Let the GPU verify caller turn on the memory probe ( #48001 ) by @tarekziade in [ #48001 ] Fix MTP config when mlp_layer_types is absent ( #48015 ) by @Cyrilvallez in [ #48015 ] Remove duplicate block_sparse_moe assignment in GraniteMoeDecoderLayer ( #47876 ) by @Aman2394 in [ #47876 ] [ModernVBERT] Fix integration test checkpoint (404 since April) ( #48009 ) by @ydshieh in [ #48009 ] Fix cropping ( #48006 ) by @Cyrilvallez in [ #48006 ] Fix DFlash candidate token device mismatch with device_map="auto" ( #47877 ) by @sywangyi in [ #47877 ] 🔴 Allow tokenizers 0.23.1 ( #46381 ) by @ArthurZucker in [ #46381 ] [Whisper] Fix batch decode_with_timestamps in WhisperTokenizer.decode() ( #47997 ) by @ydshieh in [ #47997 ] [OLMoE] Update expected logits for A10G and add torch.no_grad() ( #47989 ) by @ydshieh in [ #47989 ] [OLMo] Fix OOM in logits tests by adding torch.no_grad() ( #47986 ) by @ydshieh in [ #47986 ] [AXK1] Fix expected logits for CUDA A10G ( #47980 ) by @ydshieh in [ #47980 ] [Gemma] Update expected values for A10G ( #47976 ) by @ydshieh in [ #47976 ] Fix gemma4 video to device ( #47896 ) by @guarin in [ #47896 ] Remove stale (None, None) fallback in qwen2_5_vl batch_different_resolutions test ( #47972 ) by @ydshieh in [ #47972 ] Potential fix for code scanning alert no. 267: Artifact poisoning ( #47949 ) by @tarekziade in [ #47949 ] Fix GatedDeltaNet A_log dtype to prevent -inf under bfloat16 init ( #47944 ) by @Nkluge-correa in [ #47944 ] [serge] Fix 2 integration tests for model got_ocr2 failing with other (other (2)) ( #47937 ) by @sergereview[bot] in [ #47937 ] Fix Jinja block endings in CHAT WITH MODELS' Writing a chat template … ( #47960 ) by @ak1for2business-prog in [ #47960 ] [tests] Fix expected output for Qwen2.5-VL batch_wo_image on CUDA ( #47968 ) by @ydshieh in [ #47968 ] [serge] Fix 2 integration tests for model opt failing with other (other (2)) ( #47909 ) by @sergereview[bot] in [ #47909 ] Scan a diff in trufflehog, not the whole repo history ( #47945 ) by @tarekziade in [ #47945 ] [serge] Fix 2 integration tests for model vivit failing with output_mismatch (tensor values differ (2)) ( #47566 ) by @sergereview[bot] in [ #47566 ] Fix Gemma sliding_window being halved on every config save/reload ( #47940 ) by @Bluear7878 in [ #47940 ] docs: fix incorrect PEFT anchor link in fine-tuning section ( #47927 ) by @dsulot in [ #47927 ] CI: add vllm-test-init and vllm-test-transformers jobs on dedicated runners ( #47934 ) by @ydshieh in [ #47934 ] Make muse glimmer exportable ( #47871 ) by @IlyasMoutawwakil in [ #47871 ] docs: add installation instructions for NVIDIA Spark (ARM64) devices ( #47906 ) by @mfuntowicz in [ #47906 ] [CohereCompass] Minor docs fixes ( #47903 ) by @calpt in [ #47903 ] fix: correct checkpoints, config annotations, and create_dummy_models improvements ( #47902 ) by @ydshieh in [ #47902 ] Transform paths and repeat joining for response parsing ( #47648 ) by @Rocketknight1 in [ #47648 ] [docs] Update toctree ( #47781 ) by @stevhliu in [ #47781 ] Add CI_CPU_MEMORY_LIMIT_GB to check_failed_tests workflow ( #47884 ) by @ydshieh in [ #47884 ] Use tiny Hub checkpoint in Qwen3ASR processor test ( #47833 ) by @ydshieh in [ #47833 ] Update AutoRound XPU/CPU backend ( #47826 ) by @yiliu30 in [ #47826 ] docs(tests): fix typos in test comments ( #47859 ) by @zhaoxinyi02 in [ #47859 ] Update version post release ( #47870 ) by @Cyrilvallez in [ #47870 ] Significant community contributions The following contributors have made significant changes to the library over the last release: @tarekziade CI: gate the hunyuan-moe slow test ( #48330 ) CI: fix muse OOMs ( #48284 ) ignore mlinter ci file ( #48267 ) Post two CI badges on a PR: CPU PR CI and GPU run-slow ( #48190 ) Assign a reviewer even when a codeowner has left, and route models by modality ( #48085 ) Fix gpt_oss runs on GPU ( #48118 ) Moving mlinter to 0.1.4 ( #47918 ) unpin pytest in the examples_torch deps ( #48023 ) Let the GPU verify caller turn on the memory probe ( #48001 ) Potential fix for code scanning alert no. 267: Artifact poisoning ( #47949 ) Scan a diff in trufflehog, not the whole repo history ( #47945 ) @dkrisman Add an opt-in per-frame pixel cap (cap_pixels_per_frame) to the Qwen3-VL video processor ( #48071 ) @jiqing-feng Cpmant fix use cache ( #48013 ) [xcodec2] Fix flex attention and flash dispatch tests ( #48244 ) Always tie embeddings for LongT5 and Pop2Piano ( #47620 ) Fix CpmAnt loading: size lm_head to vocab_size ( #48012 ) Fix Qwen2.5-Omni / Qwen3-Omni-MoE generation with a compilable cache ( #47872 ) Bump default flash-attn2 hub kernel version to v3 ( #47863 ) Declare sdpa support in TimmWrapper ( #47939 ) @ydshieh Fix AutoTokenizer returning TokenizersBackend for DeepSeek-R1-Distill-Qwen models ( #48211 ) [Gemma4] Investigate flaky test_generation_beyond_sliding_window_1_eager ( #48236 ) [Gemma4] Fix stale expected values in integration tests ( #48233 ) [TableTransformer, PI0] Fix stale expected values and OOM in integration tests ( #48198 ) Fix stale expected values in integration tests (cuda sm_86 / Aug04 regressions) ( #48171 ) [EsmFold2] Fix stale expected distogram logit values ( #48182 ) Fix Apr 05 integration test regressions (cuda sm_86) ( #48170 ) Fix integration test expected values for cuda sm_86 (Mar 15 regressions) ( #48168 ) [LLaVA] Fix pixtral integration tests for cuda sm_86 ( #48166 ) [Qwen2.5-Omni] Update stale expected values for cuda sm_86 ( #48164 ) [Mistral3] Fix batched integration tests: padding_side=left + update expected values ( #48161 ) [InternVL] Fix stale expected values for Llama integration tests (cuda sm_80) ( #48153 ) Retry transient network errors (RemoteDisconnected) in github_utils ( #48124 ) [Whisper] Fix speculative decoding: preserve cleared suppress tokens through super().generate() ( #48108 ) Enable mlinter findings artifact for inline PR reviews ( #48117 ) Delete old mlinter review comments before posting new ones ( #48107 ) Accept artifact dir as argument in post_mlinter_review.py ( #48106 ) Fix mlinter artifact path ( #48088 ) [Whisper] Fix decoder position IDs for left-padded batches in longform generation ( #48028 ) [PE] Skip test_sdpa_can_dispatch_on_flash for TimmWrapper-backed models ( #48064 ) [Gemma3] Update integration test expected values for A10G ( #48036 ) [CircleCI] Enable CI for private forks, no-op for public repo ( #48056 ) [Gemma3n] Update integration test expected values for A10G + torch 2.13 ( #48035 ) [Florence2] Fix two integration test failures caused by torch 2.13 and auto-dtype ( #48031 ) [ModernVBERT] Fix integration test checkpoint (404 since April) ( #48009 ) [Whisper] Fix speculative decoding: UnboundLocalError, cache corruption, and speed regression ( #48000 ) [Whisper] Fix batch decode_with_timestamps in WhisperTokenizer.decode() ( #47997 ) [Whisper] Fix integration test failures on A10G (dtype, stale values, API changes) ( #47995 ) [DeepSeekV2] Fix integration tests OOM: use device_map=auto instead of 8-bit quantization ( #47991 ) [OLMoE] Update expected logits for A10G and add torch.no_grad() ( #47989 ) [OLMo] Fix OOM in logits tests by adding torch.no_grad() ( #47986 ) [GPTNeoX] Fix post_processor not overridden when loading from pretrained (OLMo garbage generation) ( #47988 ) [AXK1] Fix expected logits for CUDA A10G ( #47980 ) [Gemma] Update expected values for A10G ( #47976 ) [emu3] 🦮 Black Labrador is back! Fix image generation broken since #37033 ( #47948 ) Remove stale (None, None) fallback in qwen2_5_vl batch_different_resolutions test ( #47972 ) [tests] Fix expected output for Qwen2.5-VL batch_wo_image on CUDA ( #47968 ) CI: add vllm-test-init and vllm-test-transformers jobs on dedicated runners ( #47934 ) fix: correct checkpoints, config annotations, and create_dummy_models improvements ( #47902 ) Add CI_CPU_MEMORY_LIMIT_GB to check_failed_tests workflow ( #47884 ) Use tiny Hub checkpoint in Qwen3ASR processor test ( #47833 ) @eustlb gs ( #48288 ) @eladsegal Fix DynamicCache reconstruction during ExecuTorch export ( #47900 ) Support per-layer cache configuration and attention-mask selection ( #47901 ) @YangKai0616 🚨[wav2vec2] Support attn_implementation=sdpa dispatch ( #46196 ) @drbh feat: add nvfp4 quantization ( #47883 ) @Priyans-Lathiya docs: use relative paths for README language menus and add fa/ro entries ( #47777 ) @itazap [new model] step 3.7 ( #46658 ) @calpt [CohereCompass] Minor docs fixes ( #47903 ) Add CohereCompass modeling ( #47878 )

Read more →

Release v5.16.1

Release v5.16.1 This is a special release as we include GLM! (and a few small fixes) GLM-5.3-Flash GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute. Links: Documentation [Glm 5.3 Flash] GLM 5.3 Flash Support ( #48342 ) by @Dovis01 in #48342 Small patch fixes Mainly BC behavior for TP and pinning a hf kernel for security reasons 🤗 Restore BC for the tensor-parallel API ( #48300 ) by @ArthurZucker Fix kernel commit and repo paths for ESMFold2 ( #48186 ) by @Rocketknight1 Full Changelog : v5.16.0...v5.16.1

Read more →

Patch release: v5.15.1

Patch release v5.15.1 This patch most notably solves a few issues with DFlash and MTP candidate generators, as well as an issue where images could sometimes not be processed on accelerator if using Lanczos filter. It contains the following commits: Fix DFlash candidate token device mismatch with device_map="auto" ( #47877 ) by @sywangyi and @Cyrilvallez Align logit distributions for CandidateGenerators using sampling ( #48007 ) by @Cyrilvallez Fix MTP config when mlp_layer_types is absent ( #48015 ) by @Cyrilvallez Fallback from 'lanczos' to 'bicubic' when on cuda ( #48026 ) by @zucchini-nlp Fix gemma4 video to device ( #47896 ) by @guarin

Read more →

Release: v5.15.0

Release v5.15.0 New Model additions Meta Muse Glimmer Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for agentic use cases. Distilled from Muse to 30B parameters, and released under the Apache 2.0 license, it can be deployed to local setups for privacy-aware applications such as coding, document analysis, personal assistants, Claw- or Hermes-like setups. Muse Glimmer is a dense 30B parameter model consisting of: 2B ViT-style encoder for vision (Perception Encoder) 28B parameter text decoder We're covering it in the following blogpost: http://hf.co/blog/muse-glimmer GraniteMoeSWA & GraniteSWA Links: Documentation Add Granite-swa and Granitemoe-swa model support ( #47179 ) by @daviswer in #47179 Links: Documentation Add Granite-swa and Granitemoe-swa model support ( #47179 ) by @daviswer in #47179 A.X-K1 & A.X-K2 Links: Documentation Add AXK2 from SKT ( #47528 ) by @vasqu in #47528 Links: Documentation add_axk1 ( #46867 ) by @kmswin1 in #46867 Cosmos3 Edge Links: Documentation Add Cosmos3 Edge model support ( #47181 ) by @atharvajoshi10 in #47181 Breaking changes Kernels are now opt-in rather than mandatory for linear attention models (Mamba, GDN, Conv-only, etc.), so users who relied on automatic kernel selection must explicitly enable kernels to maintain previous behavior. 🚨 [ Kernels ] Refactor all linear attn models & native kernels fallback ( #47630 ) by @vasqu The cache cropping API now only accepts negative values (relative offsets) instead of absolute sizes, so users calling crop methods directly must update their code to pass negative values accordingly. 🚨 [cache] Cropping can only be done with negative values ( #47720 ) by @Cyrilvallez T5 and its model family (MT5, LongT5, etc.) now support SDPA and other attention backends via ALL_ATTENTION_FUNCTIONS , meaning the default attention implementation may change and users relying on the previous eager-only path should explicitly set attn_implementation="eager" if needed. 🚨 Enable SDPA (and other attention backends) for T5 and propagate to the T5 family ( #47014 ) by @jiqing-feng Several small private helper functions (e.g., _is_url , _build_image_tokens ) have been removed from multimodal processor files, so users or downstream libraries that imported these private functions directly must remove or replace those references. 🚨 Processors update the rest ( #46556 ) by @zucchini-nlp Attention This release includes several attention fixes and improvements, including correcting Multi-Head Latent Attention (MLA) cache compression, optimizing Flash Attention max sequence length computation in vision models, and fixing bugs in CTRL flex-attention and SDPA prefill with position bias. Additional changes refactor linear attention models for better maintainability, make Gemma 4's heterogeneous attention config explicit, and improve MPS support via metal-flash-sdpa integration. [Fix] Fix multi-head latent attention (MLA) ( #47761 ) by @remi-or in [ #47761 ] Refactor all linear attention models to latest best standards for convolution ( #47452 ) by @Cyrilvallez in [ #47452 ] Allow metal-flash-sdpa for OpenAIPrivacyFilter on MPS ( #46740 ) by @ArthurZucker in [ #46740 ] Use new per_layer_config for Gemma 4 so that heterogeneous attention config is explicit ( #47384 ) by @hmellor in [ #47384 ] add paged attention tests support for XPU ( #47163 ) by @kaixuanliu in [ #47163 ] Move value padding into the attention interfaces that need it ( #47451 ) by @hmellor in [ #47451 ] Simplify function dispatch for linear attention ( #47450 ) by @Cyrilvallez in [ #47450 ] Optimize flash attention max seqlen computation in vision attention ( #47170 ) by @ShareLer in [ #47170 ] Fix BlockMask crash in CTRL flex-attention generation ( #46854 ) by @jiqing-feng in [ #46854 ] [CB] Automatically switch attention implementation to flash ( #47330 ) by @remi-or in [ #47330 ] Fix sdpa prefill with position_bias ( #47359 ) by @Cyrilvallez in [ #47359 ] Vision Vision improvements in this release include performance optimizations such as faster image preprocessing for vision-language models (GLM4V, MiniMaxM3-VL, and others) by eliminating redundant tensor copies, and more efficient Flash Attention variable-length paths by precomputing maximum sequence lengths once per forward pass. Several bug fixes were also applied, including correcting dtype alignment in Kosmos2/Kosmos2_5 embedding merges, fixing a position-embedding initialization fallback in Phi4Multimodal, resolving PIL resize parity in Hunyuan-VL, and patching stop-sequence handling in the image-text-to-text pipeline. Modularize qwen-format vision processors ( #47573 ) by @zucchini-nlp in [ #47573 ] Update daily CI Docker image to torch 2.13.0 / CUDA 13.0 ( #47738 ) by @ydshieh in [ #47738 ] Align image feature dtype in kosmos2 and kosmos2_5 embedding merge ( #47691 ) by @ in [ #47691 ] Speed up image preprocessing for vision-language models ( #47453 ) by @labAxiaoming in [ #47453 ] Fix vision position-embedding init width fallback in Phi4Multimodal ( #47509 ) by @ in [ #47509 ] Fix Hunyuan-VL PIL image resize parity with reference preprocessing ( #47233 ) by @IMvision12 in [ #47233 ] Fix image-text-to-text stop_sequence handling ( #47032 ) by @Sunt-ing in [ #47032 ] Refactor image loading in tests to use load_test_image helper ( #47218 ) by @LevelVoid in [ #47218 ] Generation Several generation improvements and bug fixes were made, including enabling batched audio generation for Qwen2.5/3-Omni, allowing sliding window cache layers to work with speculative decoding, and fixing memory overhead from static cache persistence across generate() calls. Multiple model-specific bugs were also resolved, including crashes in KyutaiSpeechToText, MusicgenForCausalLM, CTRL flex-attention, and assisted decoding for EncoderDecoder cache and OlmoHybrid models. Align OlmoHybrid to use a native cache in generate ( #47604 ) by @Cyrilvallez in [ #47604 ] [generate] Stop setting the static cache as an attribute to save memory ( #47731 ) by @Cyrilvallez in [ #47731 ] Add support for batched Qwen2.5/3-Omni audio generation ( #47186 ) by @IMvision12 in [ #47186 ] [cache] Allow sliding window layers to be roll-backed for speculative decoding ( #47447 ) by @Cyrilvallez in [ #47447 ] Fix shape mismatch in KyutaiSpeechToText generate() last window ( #46952 ) by @jiqing-feng in [ #46952 ] Fix typo in MusicgenForCausalLM.generate() ( #46974 ) by @jiqing-feng in [ #46974 ] Fix assisted decoding for models with EncoderDecoder cache & OlmoHybrid ( #47361 ) by @Cyrilvallez in [ #47361 ] Cache Several cache-related bugs were fixed, including correcting NemotronH's missing "mlp" layer-type mapping, resolving recurrent-layer padding masks being skipped during chunked prefill and cache continuation for hybrid models, and fixing assisted decoding for models with EncoderDecoderCache and OlmoHybrid. Additional improvements include aligning OlmoHybrid to use a native cache, enabling sliding window layers to support speculative decoding rollback, and stopping the static cache from being stored as a model attribute to reduce unexpected memory overhead. [docs] MPS graph cache ( #47304 ) by @stevhliu in [ #47304 ] Fix NemotronH: Register "mlp" in the cache layer-type mappings ( #47535 ) by @qgallouedec in [ #47535 ] Fix recurrent-layer padding mask being skipped on continued forwards (chunked prefill, cache continuation) ( #47087 ) by @abcgco in [ #47087 ] Kernels ⚠️ The kernels python package will very likely be a required dependency for transformers[torch] in the near future. This will help us deliver maximum performance to all users; kernels will only be downloaded from trusted publishers manually approved by the HF team. Please let us know of any issues you're facing beforehands so that we may solidify our integration. Improved robustness of the kernels integration by refactoring function handling to use layer repos, fixing CI EROFS fallback patches for kernel downloads via HfApi , resolving a positional argument collision in causal_conv1d_fn , and bumping the FP8 kernels version to prevent NaNs. [conftest] Fix EROFS fallback for kernel downloads (correct interception point) ( #47794 ) by @ydshieh in [ #47794 ] [conftest] Fix EROFS fallback for kernel downloads via HfApi ( #47791 ) by @ydshieh in [ #47791 ] [ Kernels ] Refactor function handling ( #46883 ) by @vasqu in [ #46883 ] Kernels and loaders robustification ( #47334 ) by @IlyasMoutawwakil in [ #47334 ] Fix causal_conv1d_fn positional activation colliding with hub kernel's seq_idx ( #47527 ) by @qgallouedec in [ #47527 ] [ FP8 ] Bump kernels version ( #47344 ) by @vasqu in [ #47344 ] [docs] FlashAttention kernel fallback ( #47345 ) by @stevhliu in [ #47345 ] Quantization Quantization support was expanded with FP8 kernels for compressed-tensors models, fixes for FP8 module normalization and format-based compression detection, and a multi-device MXFP4 dequantization race condition fix. GPTQ and MXFP4 tests were also extended to cover Intel XPU devices. extend tests/quantization/gptq/test_gptq.py::GPTQTestCUDA and tests/q… ( #47166 ) by @sywangyi in [ #47166 ] Compressed tensors fp8 ( #47216 ) by @SunMarc in [ #47216 ] Fix A.X-K2 fp8 modules_to_not_convert normalization for the gated-norm MLP ( #47578 ) by @kmswin1 in [ #47578 ] [Quantization]: Refactor is_quantization_compressed for format-based detection ( #47152 ) by @rigen1048 in [ #47152 ] Fix multi-device mxfp4 dequantization race in _convert_moe_packed_tensors ( #47423 ) by @kaixuanliu in [ #47423 ] Audio Batched audio generation is now supported for Qwen2.5/3-Omni, and several bug fixes were applied across audio models, including a dtype mismatch in Gemma4 audio feature merging, a bfloat16 positional embedding error in AudioFlamingo3, and missing backend requirement guards for Voxtral. The VibeVoice ASR processor was also updated to make audio input optional and support multiple audios per prompt. feat[vLLM x v5]: Make audio optional and support multiple audios in VibeVoice ASR processor ( #47483 ) by @harshaljanjani in [ #47483 ] Fix Gemma4 audio feature dtype mismatch in masked_scatter ( #47482 ) by @danielhanchen in [ #47482 ] [fix] fix requirements audio feature and proc ( #47113 ) by @eustlb in [ #47113 ] [AudioFlamingo3] Fix bfloat16 dtype mismatch in audio encoder positional embedding ( #47258 ) by @snkii in [ #47258 ] Parallelization Expanded FSDP support across 94 ForCausalLM model classes with auto-generated FSDP plans, added end-to-end FSDP tests including distributed checkpoint save/load and generation, and introduced a dedicated FSDP CI job. Additionally, fixed a device mismatch bug in create_bidirectional_sliding_window_mask under model parallelism and resolved a tensor parallel inference issue for models with tied embeddings. skip fsdp tests when backend is mps ( #47601 ) by @3outeille in [ #47601 ] Fix model parallel device mismatch in create_bidirectional_sliding_window_mask ( #47560 ) by @abcgco in [ #47560 ] Add FSDP plans to all models ( #47165 ) by @3outeille in [ #47165 ] Fix TP inference for tied embedding ( #47503 ) by @3outeille in [ #47503 ] Add FSDP CI and end-to-end FSDP tests + save fsdp ( #47357 ) by @3outeille in [ #47357 ] Tokenization This release adds native support for Mistral's "tekken" tokenizer format via AutoTokenizer, fixes a CodeLlama tokenizer bug where leading whitespace was incorrectly dropped during decode, and patches a potential ReDoS vulnerability caused by unescaped tokenizer filenames being used as regex patterns in from_pretrained . [Mistral] Add native tekken tokenizer support to AutoTokenizer ( #47507 ) by @juliendenize in [ #47507 ] Fix CodeLlama tokenizer dropping leading whitespace on decode ( #47488 ) by @SuryanshSS1011 in [ #47488 ] Fix potential ReDoS by escaping tokenizer filename used as regex pattern ( #47498 ) by @hameedibrh in [ #47498 ] Serve Improved the serve chat parsing to unify streaming and non-streaming paths under a single response parser that handles tool calls, reasoning, and content, simplifying the addition of new model support. Additionally, hardened daily CI reporting by fixing GitHub API diagnostic output being captured in Slack payloads and adding rate-limit resilience to prevent report failures when paginating large job matrices. CI: Log GitHub API diagnostics to stderr ( #47635 ) by @tarekziade in [ #47635 ] Update serve chat parsing ( #46267 ) by @SunMarc in [ #46267 ] ci: harden daily CI reporting against GitHub API rate limits ( #47382 ) by @tarekziade in [ #47382 ] Bugfixes and improvements Fix cached_files silently returning stale file on read-only filesystem (EROFS) ( #47852 ) by @ydshieh in [ #47852 ] Fix PhimoeIntegrationTest ( #46539 ) by @ydshieh in [ #46539 ] make examples under doc device agnostic ( #47812 ) by @kaixuanliu in [ #47812 ] cancel deterministic for XPU in gemma4 tests ( #47790 ) by @kaixuanliu in [ #47790 ] Serialize post-mlinter-review after post-link to avoid PR description race ( #47832 ) by @ydshieh in [ #47832 ] Use content hash for mlinter review deduplication ( #47830 ) by @ydshieh in [ #47830 ] Add new args in auto-docstring ( #47737 ) by @zucchini-nlp in [ #47737 ] Add check_model_inits.py ( #47656 ) by @guarin in [ #47656 ] add xpu in installation guide ( #47785 ) by @sywangyi in [ #47785 ] Fix mlinter review job: checkout before artifact download ( #47820 ) by @ydshieh in [ #47820 ] Post mlinter findings as inline PR review comments ( #47819 ) by @ydshieh in [ #47819 ] Fix ci style ( #47818 ) by @vasqu in [ #47818 ] Hotfix axk2 indexer norm ( #47810 ) by @kmswin1 in [ #47810 ] open fla support for XPU to benefit from the acceleration ( #47799 ) by @kaixuanliu in [ #47799 ] [docs] Update BatchEncoding.to() type annotation and docstring ( #47789 ) by @samyuktahegde in [ #47789 ] Fix linting ( #47807 ) by @Cyrilvallez in [ #47807 ] Fix patching in some models ( #47798 ) by @zucchini-nlp in [ #47798 ] Add post-mlinter-review job to post-dashboard-link workflow ( #47800 ) by @ydshieh in [ #47800 ] Migrate torchao integration off deleted torchao.dtypes ( #47797 ) by @vkuzo in [ #47797 ] [conftest] Also wrap snapshot_download for EROFS fallback ( #47796 ) by @ydshieh in [ #47796 ] Fix MI355 CI: bump hf-workflows pin to NUM_SLICES=4 ( #47792 ) by @Abdennacer-Badaoui in [ #47792 ] [Fix] Wrong type hint in get_number_of_image_patches ( #47788 ) by @remi-or in [ #47788 ] Fix Dac offload tests ( #47775 ) by @guarin in [ #47775 ] [Fix] Swapped height and width in KimiK25 ( #47786 ) by @remi-or in [ #47786 ] Remove dangling files and folders ( #47764 ) by @Cyrilvallez in [ #47764 ] Fix AI-written conversion mappings ( #47755 ) by @Cyrilvallez in [ #47755 ] update mistral common version for PR 47507 ( #47677 ) by @itazap in [ #47677 ] Fix spelling/grammar in model files (batch 3/3) ( #47684 ) by @Rocketknight1 in [ #47684 ] Fix spelling/grammar in core library, examples, and utils ( #47685 ) by @Rocketknight1 in [ #47685 ] Fix spelling/grammar in model files (batch 2/3) ( #47683 ) by @Rocketknight1 in [ #47683 ] update sonicmoe versions ( #47769 ) by @IlyasMoutawwakil in [ #47769 ] clean up reverse_op fixme in compressed_tensors ( #47701 ) by @DhanushPillay in [ #47701 ] PR CI with torch 2.13 ( #47767 ) by @ydshieh in [ #47767 ] Import utils compilation fixes ( #47726 ) by @IlyasMoutawwakil in [ #47726 ] Fix DBRX MoE hidden size and expert GLU transposes ( #47671 ) by @kaixuanliu in [ #47671 ] Fix multi token decode merging ( #47762 ) by @IlyasMoutawwakil in [ #47762 ] Fix: Remove redundant @can_return_tuple conflicting with @capture_out… ( #47733 ) by @guarin in [ #47733 ] fix processor config nested key fallback ( #47628 ) by @YunzhuLu in [ #47628 ] [CI - Debug] Skip /transformers-dependent steps for CPU runner ( #47759 ) by @ydshieh in [ #47759 ] [CI] Add CPU runner support to ssh-runner workflow ( #47757 ) by @ydshieh in [ #47757 ] Simplify reverse weight conversion ( #47725 ) by @Cyrilvallez in [ #47725 ] Remove useless linting for inv_freq ( #47753 ) by @Cyrilvallez in [ #47753 ] Use explicit nn.Buffer everywhere for modular ( #47722 ) by @Cyrilvallez in [ #47722 ] Remove stale and redundant _no_split_modules entries ( #47645 ) by @guarin in [ #47645 ] Executorch exporter fixes ( #47243 ) by @IlyasMoutawwakil in [ #47243 ] Acc fix in xpu ( #47500 ) by @sywangyi in [ #47500 ] update Dockerfile for xpu torch2.13 ( #47502 ) by @sywangyi in [ #47502 ] feat[vLLM]: Support text replacement offsets in the remaining old-format processors ( #47614 ) by @harshaljanjani in [ #47614 ] Fix ImportError in transformers.exporters on torch < 2.8 ( #47711 ) by @Neal006 in [ #47711 ] Fix failing tests for axk1 and axk2 ( #47727 ) by @kaixuanliu in [ #47727 ] Fix failing tests for granite_swa and granitemoe_swa ( #47723 ) by @kaixuanliu in [ #47723 ] skip invalid test cases for inkling tests ( #47493 ) by @kaixuanliu in [ #47493 ] [docs] Fix BatchEncoding documentation inconsistencies ( #47647 ) by @samyuktahegde in [ #47647 ] Fix spelling/grammar in model files (batch 1/3) ( #47682 ) by @Rocketknight1 in [ #47682 ] Fix spelling/grammar in English docs (q → z) ( #47681 ) by @Rocketknight1 in [ #47681 ] Fix spelling/grammar in English docs (h → p) ( #47680 ) by @Rocketknight1 in [ #47680 ] Fix spelling/grammar in English docs (a → g) ( #47679 ) by @Rocketknight1 in [ #47679 ] fix npu check ( #47587 ) by @DhanushPillay in [ #47587 ] fix: correct text input validation logic in 8 multimodal processors (and → or) ( #47663 ) by @AbdullahRasheed45 in [ #47663 ] Fix feature dtype mismatch in masked_scatter for seven multimodal models ( #47673 ) by @ in [ #47673 ] [docs] storing and loading chat templates ( #47650 ) by @stevhliu in [ #47650 ] adding amd quark config class changes ( #47322 ) by @debasisdwivedy in [ #47322 ] Fix compressed tensors impl ( #47652 ) by @SunMarc in [ #47652 ] Hoist special-token lookups in wav2vec2 decode paths and drop a dead filter in wav2vec2_phoneme ( #47557 ) by @ishan-1010 in [ #47557 ] silencing elastic warning by import distributed lib inside functions ( #47665 ) by @3outeille in [ #47665 ] Fix missing github_utils.py download in PR CI dashboard workflow ( #47668 ) by @ydshieh in [ #47668 ] Exportable kimi ( #47096 ) by @IlyasMoutawwakil in [ #47096 ] better guarding to handle torch compiled with USE_DISTRIBUTED=0 ( #47619 ) by @3outeille in [ #47619 ] Remove gemma4 warnings ( #47664 ) by @Cyrilvallez in [ #47664 ] [Chat Parsing] Type inline tool-call arguments from the calling tool's JSON Schema ( #47529 ) by @yonigozlan in [ #47529 ] Improve Trainer DataLoader Controls for Streaming and Multiprocessing ( #47164 ) by @muyihao in [ #47164 ] Update maintainer list ( #47644 ) by @Rocketknight1 in [ #47644 ] Allow position_ids_start=2 on DataCollatorWithFlattening for RoBERTa etc. ( #47525 ) by @tomaarsen in [ #47525 ] Remove Rotary warning ( #47642 ) by @Cyrilvallez in [ #47642 ] Drop multimodal inputs natively in prepare_inputs_for_generation if not in prefill ( #47622 ) by @Cyrilvallez in [ #47622 ] Simplify all Rotary modules ( #47598 ) by @Cyrilvallez in [ #47598 ] [docs] response_template when serving ( #47626 ) by @stevhliu in [ #47626 ] [docs] MTP support ( #47301 ) by @stevhliu in [ #47301 ] Fix GPT-2 c_proj depth scaling initialization ( #47459 ) by @DavidJohnQuinlan in [ #47459 ] Fix fp8_linear compilability ( #47623 ) by @IlyasMoutawwakil in [ #47623 ] byebye torch 2.4 ( #47609 ) by @ydshieh in [ #47609 ] Vectorize NoRepeatNGramLogitsProcessor and remove its host sync ( #47571 ) by @hameedibrh in [ #47571 ] Fix CUDA Graph breaking host to device copy from scalar tensor allocation ( #47547 ) by @hmellor in [ #47547 ] Fix some processors ( #47608 ) by @zucchini-nlp in [ #47608 ] CI: Add serge review relay workflow and review rules ( #47610 ) by @tarekziade in [ #47610 ] [docs] Exporters ( #47374 ) by @stevhliu in [ #47374 ] CI: use a single function for GH calls ( #47474 ) by @tarekziade in [ #47474 ] Remove redundant guarding for distributed ( #47570 ) by @3outeille in [ #47570 ] Fix modular for mamba packages ( #47494 ) by @Cyrilvallez in [ #47494 ] Remove deprecated conversion in Kimi ( #47581 ) by @zucchini-nlp in [ #47581 ] CI: use transformers-ci daily workflow with OTEL ( #47360 ) by @tarekziade in [ #47360 ] Fix slow tensor path in _check_special_mm_tokens ( #47580 ) by @guan404ming in [ #47580 ] CI: let's run integration failure cron at 10pm ( #47537 ) by @tarekziade in [ #47537 ] Better and more extensive tests for RoPE ( #46912 ) by @zucchini-nlp in [ #46912 ] Fix mamba2 family decode and simplify all reshape ops ( #47569 ) by @Cyrilvallez in [ #47569 ] [DiffusionGemma] Cast the decoder padding mask to bool ( #47295 ) by @kashif in [ #47295 ] General maintenance ( #47517 ) by @zucchini-nlp in [ #47517 ] fixed the benchmark script with DistributedConfig ( #47568 ) by @tarekziade in [ #47568 ] Fix failing tests for mimo_v2_flash ( #47284 ) by @kaixuanliu in [ #47284 ] Make tokenization_mistral_common importable without mistral_common installed ( #47397 ) by @juliendenize in [ #47397 ] Use --flake-runs=1 in check_bad_commit.py for PR comment CI ( #47522 ) by @ydshieh in [ #47522 ] Deprecate the old response_schema ( #47320 ) by @Rocketknight1 in [ #47320 ] fix: add pickle support to _LazyConfigMapping for spawn multiprocessing ( #46026 ) by @kfojcik-intel in [ #46026 ] Delete old deprecations ( #47518 ) by @zucchini-nlp in [ #47518 ] Run only @slow tests in PR comment CI ( #47521 ) by @ydshieh in [ #47521 ] Fix incorrect type hint ( #47519 ) by @hmellor in [ #47519 ] Deprecate CB config in gen configuration ( #47291 ) by @remi-or in [ #47291 ] [Offloading] [Bugfix] Fix fully offloaded model saving ( #47336 ) by @kylesayrs in [ #47336 ] CI: fix torchaudio pinning +proper break in rnnt ( #47422 ) by @tarekziade in [ #47422 ] Fix Qwen2.5-Omni Token2Wav DiT rotary embedding layout (interleaved cos/sin) ( #47403 ) by @HenryVarro666 in [ #47403 ] CI: add reproduce mode to serge verify caller ( #47492 ) by @tarekziade in [ #47492 ] [fix][whisper]: fix max_new_tokens handling ( #46795 ) by @eustlb in [ #46795 ] Isolate MLA KV expansion to make it easier to bypass ( #47460 ) by @hmellor in [ #47460 ] Fix loss alignment and Trainer token counting for encoder decoder models ( #46903 ) by @OmkumarSolanki in [ #46903 ] Fix the HunyuanVL's torchvision backend ( #47499 ) by @Mi-Jiazhi in [ #47499 ] [DiffusionGemma] Support gradient checkpointing ( #46572 ) by @kashif in [ #46572 ] Fix failing tests for zaya ( #47268 ) by @kaixuanliu in [ #47268 ] tipsv2_dpt: fix failing tests for XPU ( #47292 ) by @kaixuanliu in [ #47292 ] Fix some failed test cases related with XPU Expectations ( #47173 ) by @kaixuanliu in [ #47173 ] fix: guard DTensor import in sharding_utils.py for PyTorch < 2.5 ( #47481 ) by @ in [ #47481 ] CPU can incur a slow path on non-contiguous magnitudes ( #47351 ) by @vbayanag in [ #47351 ] Route chat management calls to the service root ( #47138 ) ( #47303 ) by @dhruv7477 in [ #47303 ] fix: liger unnecessarily materializes logits in VRAM during eval, causing OOM ( #45273 ) by @excepshenal in [ #45273 ] Fix gradient inflation when combining label smoothing with gradient accumulation ( #47261 ) by @Incheonkirin in [ #47261 ] Fix MoE expert decompression for non-32-divisible bit widths ( #47315 ) by @KKothuri in [ #47315 ] [Qwen3ASR] Add hotword parsing, and fix language parsing and training. ( #47111 ) by @ebezzam in [ #47111 ] Remove deprecated training args and is_fast property ( #46917 ) by @cyyever in [ #46917 ] Consistent output shape from get_image_features ( #46405 ) by @zucchini-nlp in [ #46405 ] fix failed test cases for qwen3_omni_moe model ( #47449 ) by @kaixuanliu in [ #47449 ] Fix double-shifted training loss in GitForCausalLM ( #47395 ) by @ in [ #47395 ] Fix CohereASR training-loss double-shift (same as Moonshine fix #46784 ) ( #46895 ) by @sharmax-vikas in [ #46895 ] Warn when group_by_length is silently ignored for iterable datasets ( #47379 ) by @qgallouedec in [ #47379 ] Update bug report list ( #46607 ) by @molbap in [ #46607 ] fix: remove unreachable return in special token builder ( #47420 ) by @hai1222 in [ #47420 ] Add Harry to slow CI ( #47454 ) by @vasqu in [ #47454 ] BLT: vectorize patch length processing ( #47385 ) by @sj0618 in [ #47385 ] Fix TrackioCallback fails to log evaluation metrics after training ends ( #46935 ) by @lewtun in [ #46935 ] [Kimi] add integration tests ( #47383 ) by @zucchini-nlp in [ #47383 ] Add distributed runtime utils and DistributedMixin ( #47352 ) by @3outeille in [ #47352 ] Fix Cosmos 3 Edge Patch packing order ( #47399 ) by @atharvajoshi10 in [ #47399 ] Fix yarn mscale_all_dim for DeepSeek v2 and Mistral 4 ( #47435 ) by @hmellor in [ #47435 ] Hoist special-token lookups out of per-token loops in six slow tokenizers ( #47425 ) by @ishan-1010 in [ #47425 ] fix typos and variable naming in quicktour.md ( #47418 ) by @yashasvi-srivastava21 in [ #47418 ] Normalize multimodal input keys in AnyToAnyPipeline ( #47074 ) by @Sunt-ing in [ #47074 ] Fix Aria checkpoint key conversion mapping ( #47151 ) by @sywangyi in [ #47151 ] extend tests/models/qwen3_next/test_modeling_qwen3_next.py::Qwen3Next… ( #47184 ) by @sywangyi in [ #47184 ] CI: add serge verify (GPU) caller workflow ( #47381 ) by @tarekziade in [ #47381 ] Fix orthogonal_ init for low-precision dtypes (bf16/fp16) ( #47252 ) by @janbernloehr in [ #47252 ] [docs] Fix decode examples and expected output in fast_tokenizers ( #47369 ) by @samyuktahegde in [ #47369 ] [serge] Fix 20 integration tests for model whisper failing with output_mismatch (list output differs (10), other (6) ( #47150 ) by @sergereview[bot] in [ #47150 ] [Mistral] Move MistralConverter into integrations/mistral/ package ( #46603 ) by @juliendenize in [ #46603 ] Fix GLM video frame padding for temporal patches ( #47141 ) by @labAxiaoming in [ #47141 ] [docs] Inkling ( #47350 ) by @stevhliu in [ #47350 ] Fix Daily CI reporting issues ( #47364 ) by @tarekziade in [ #47364 ] [ peft ] Support key_mapping with PEFT models ( #46766 ) by @tomaarsen in [ #46766 ] Fix model tests for tipsv2 ( #47356 ) by @kaixuanliu in [ #47356 ] Fix TimesFM 2.5 window_size AttributeError ( #47363 ) by @kashif in [ #47363 ] fix: allow num_labels property to return None when id2label is unset ( #47069 ) by @SebTardif in [ #47069 ] [Tests] Fix slow video tensor creation from list of numpy arrays in SmolVLM ( #44731 ) by @Defalt-Meh in [ #44731 ] Update dev version on main ( #47366 ) by @vasqu in [ #47366 ] Fix deepgemm on multiple devices ( #47323 ) by @IlyasMoutawwakil in [ #47323 ] ci: Handle empty GitHub token in CI run lookup ( #47362 ) by @tarekziade in [ #47362 ] Fix inkling feature extractor ( #47349 ) by @ArthurZucker in [ #47349 ] Significant community contributions The following contributors have made significant changes to the library over the last release: @ydshieh Fix cached_files silently returning stale file on read-only filesystem (EROFS) ( #47852 ) Fix PhimoeIntegrationTest ( #46539 ) Serialize post-mlinter-review after post-link to avoid PR description race ( #47832 ) Use content hash for mlinter review deduplication ( #47830 ) Fix mlinter review job: checkout before artifact download ( #47820 ) Post mlinter findings as inline PR review comments ( #47819 ) Add post-mlinter-review job to post-dashboard-link workflow ( #47800 ) [conftest] Also wrap snapshot_download for EROFS fallback ( #47796 ) [conftest] Fix EROFS fallback for kernel downloads (correct interception point) ( #47794 ) [conftest] Fix EROFS fallback for kernel downloads via HfApi ( #47791 ) PR CI with torch 2.13 ( #47767 ) [CI - Debug] Skip /transformers-dependent steps for CPU runner ( #47759 ) [CI] Add CPU runner support to ssh-runner workflow ( #47757 ) Update daily CI Docker image to torch 2.13.0 / CUDA 13.0 ( #47738 ) Fix missing github_utils.py download in PR CI dashboard workflow ( #47668 ) byebye torch 2.4 ( #47609 ) Use --flake-runs=1 in check_bad_commit.py for PR comment CI ( #47522 ) Run only @slow tests in PR comment CI ( #47521 ) @kaixuanliu make examples under doc device agnostic ( #47812 ) cancel deterministic for XPU in gemma4 tests ( #47790 ) open fla support for XPU to benefit from the acceleration ( #47799 ) Fix DBRX MoE hidden size and expert GLU transposes ( #47671 ) Fix failing tests for axk1 and axk2 ( #47727 ) Fix failing tests for granite_swa and granitemoe_swa ( #47723 ) skip invalid test cases for inkling tests ( #47493 ) Fix failing tests for mimo_v2_flash ( #47284 ) Fix failing tests for zaya ( #47268 ) tipsv2_dpt: fix failing tests for XPU ( #47292 ) Fix some failed test cases related with XPU Expectations ( #47173 ) add paged attention tests support for XPU ( #47163 ) Fix multi-device mxfp4 dequantization race in _convert_moe_packed_tensors ( #47423 ) fix failed test cases for qwen3_omni_moe model ( #47449 ) Fix model tests for tipsv2 ( #47356 ) @vasqu Fix ci style ( #47818 ) 🚨 [ Kernels ] Refactor all linear attn models & native kernels fallback ( #47630 ) [ Kernels ] Refactor function handling ( #46883 ) Add AXK2 from SKT ( #47528 ) Add Harry to slow CI ( #47454 ) Update dev version on main ( #47366 ) [ FP8 ] Bump kernels version ( #47344 ) @kmswin1 Hotfix axk2 indexer norm ( #47810 ) Fix A.X-K2 fp8 modules_to_not_convert normalization for the gated-norm MLP ( #47578 ) add_axk1 ( #46867 ) @remi-or [Fix] Fix multi-head latent attention (MLA) ( #47761 ) [Fix] Wrong type hint in get_number_of_image_patches ( #47788 ) [Fix] Swapped height and width in KimiK25 ( #47786 ) Deprecate CB config in gen configuration ( #47291 ) [CB] Automatically switch attention implementation to flash ( #47330 ) @juliendenize [Mistral] Add native tekken tokenizer support to AutoTokenizer ( #47507 ) Make tokenization_mistral_common importable without mistral_common installed ( #47397 ) [Mistral] Move MistralConverter into integrations/mistral/ package ( #46603 ) @IMvision12 Add support for batched Qwen2.5/3-Omni audio generation (#47186) Fix Hunyuan-VL PIL image resize parity with reference preprocessing (#47233) @jiqing-feng 🚨 Enable SDPA (and other attention backends) for T5 and propagate to the T5 family (#47014) Fix shape mismatch in KyutaiSpeechToText generate() last window (#46952) Fix typo in MusicgenForCausalLM.generate() (#46974) Fix BlockMask crash in CTRL flex-attention generation (#46854) @tarekziade CI: Log GitHub API diagnostics to stderr (#47635) CI: Add serge review relay workflow and review rules (#47610) CI: use a single function for GH calls (#47474) CI: use transformers-ci daily workflow with OTEL (#47360) CI: let's run integration failure cron at 10pm (#47537) fixed the benchmark script with DistributedConfig (#47568) CI: fix torchaudio pinning +proper break in rnnt (#47422) CI: add reproduce mode to serge verify caller (#47492) CI: add serge verify (GPU) caller workflow (#47381) ci: harden daily CI reporting against GitHub API rate limits (#47382) Fix Daily CI reporting issues (#47364) ci: Handle empty GitHub token in CI run lookup (#47362) @daviswer Add Granite-swa and Granitemoe-swa model support (#47179) @ShareLer Optimize flash attention max seqlen computation in vision attention (#47170) @atharvajoshi10 Fix Cosmos 3 Edge Patch packing order (#47399) Add Cosmos3 Edge model support (#47181)

Read more →

Patch release: v5.14.1

Patch release v5.14.1 This patch solves a few issues which appeared when integrating Inkling model, most notably an issue affecting models using EncoderDecoderCache during assisted generation. It also fixes an issue that could appear during prefill with StaticCache and sdpa without padding for Inkling which uses a position_bias. It contains the following commits: Fix sdpa prefill with position_bias ( #47359 ) by @Cyrilvallez Fix assisted decoding for models with EncoderDecoder cache & OlmoHybrid ( #47361 ) by @Cyrilvallez [FP8] Bump kernels version ( #47344 ) by @vasqu Fix deepgemm on multiple devices ( #47323 ) by @IlyasMoutawwakil

Read more →

Release v5.14.0

Release v5.14.0 New Model additions Inkling (fresh from Thinking Machines): 975B total, 41B active Add Inkling model #47347 by @molbap @Cyrilvallez @eustlb and @zucchini-nlp Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI- powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers. TIPSv2 Links: Documentation Add TIPSv2 ( #46347 ) by @Ternura143 in #46347 TIPSv2 DPT Links: Documentation Add TIPSv2 ( #46347 ) by @Ternura143 in #46347 🚨 Breaking changes GPTNeoX now remaps embed_out to lm_head and GPTBigCode has _supports_attention_backend = True enabled for vLLM compatibility; users relying on the previous weight naming or attention backend behavior for these models should update their code accordingly. 🚨 Fix GPTBigCode and GPTNeoX for the Transformers modelling backend for vLLM ( #47198 ) by @hmellor Kernels Several kernel-related fixes and improvements were made, including pinning the kernels dependency to a compatible version in the benchmark workflow, removing a deprecated package_name argument from LocalLayerRepository , and making the DeepGEMM Triton fallback more robust when CUDA_HOME is unset or misconfigured. Additionally, SDPA prefill was updated to leverage the FlashAttention kernel with StaticCache , yielding significant performance gains (up to 260% faster for large input sizes). Pin kernels to compatible version in benchmark workflow ( #47339 ) by @tarekziade in [ #47339 ] [Fix] Remove deprecated argument from kernels call ( #47100 ) by @remi-or in [ #47100 ] [Fix] Make DeepGEMM triton fallback more robust ( #47126 ) by @remi-or in [ #47126 ] [sdpa] Allow prefill to use FA kernel with StaticCache ( #47094 ) by @Cyrilvallez in [ #47094 ] Generation Generation improvements include adding Multi-Token Prediction (MTP) decoding support, static ensemble verification for speculative decoding to improve draft token acceptance rates, and a fix for crashes in greedy assisted generation with different tokenizers. A misleading double-negative warning message for synced_gpus in continuous batching mode was also corrected. [generation] Fix misleading synced_gpus warning in continuous batching ( #47158 ) by @Partha-Shankar in [ #47158 ] [generate] Add proper MTP support ( #46229 ) by @Cyrilvallez in [ #46229 ] Fix crash in greedy assisted generation with different tokenizers ( #46936 ) by @Sunt-ing in [ #46936 ] [Generation] Add static ensemble verification for lossy speculative decoding ( #45979 ) by @kasakh in [ #45979 ] Performance Fixed a Flash Attention performance regression affecting models like Qwen3-VL and resolved a MoE decode optimization bug where the grouped-to-batched matrix multiplication switch was not applied to experts residing in submodels (e.g., VLMs with a nested text config). Fix FA performance regression ( #47134 ) by @andreasgoulas in [ #47134 ] Fix MoE decode optimization for experts living in a submodel ( #47107 ) by @IlyasMoutawwakil in [ #47107 ] Make doc builds faster ( #47099 ) by @mishig25 in [ #47099 ] Cache Cache dispatch logic was simplified by introducing explicit layer-type mappings for sliding and static layers, reducing complexity in cache routing. Additionally, fixes were made for read-only cache failures in CPU CI environments and for MPS graph cache growth during variable-length batch training on Apple Silicon. Fix CI read-only cache failures by patching cached_files in conftest ( #47043 ) by @ydshieh in [ #47043 ] trainer: clear MPS graph cache via torch_empty_cache_steps ( #45818 ) by @anagnorisis2peripeteia in [ #45818 ] [cache] Simplify cache dispatch based on layer_types ( #47118 ) by @Cyrilvallez in [ #47118 ] Bugfixes and improvements ci: cover xet as well (runtime error) ( #47338 ) by @tarekziade in [ #47338 ] [docs] TokenizersBackend fallback ( #47302 ) by @stevhliu in [ #47302 ] Resolve continuous batching XPU availability checks at runtime ( #47185 ) by @kaixuanliu in [ #47185 ] [Nit] Add kernels_fallback_ok kwarg to is_flash_attn_N_available ( #47318 ) by @remi-or in [ #47318 ] [Nit] Add expectations for gemma4 tests on H100 ( #47311 ) by @remi-or in [ #47311 ] [docs] DeepGEMM requirements ( #47324 ) by @stevhliu in [ #47324 ] DeepGEMM shouldn't pad on SM90 ( #47313 ) by @IlyasMoutawwakil in [ #47313 ] Fix half-precision torch.compile crash in DETR-family sine position embeddings ( #47238 ) by @David-Wu1119 in [ #47238 ] Fix hardcoded paths in siglip checkpoint/vocab loading ( #47178 ) by @XanxusCrypto in [ #47178 ] Update AMD CI runner groups to amd-mi300 ( #47307 ) by @Abdennacer-Badaoui in [ #47307 ] Point to Gemma 4 model in Gemma4ForCausalLM docstring example ( #47255 ) by @lefft in [ #47255 ] Fix Qwen Omni batched text postprocessing ( #47197 ) by @Sunt-ing in [ #47197 ] Fix AqlmConfig error messages to say "int" instead of "float" ( #47089 ) by @Sreekant13 in [ #47089 ] Fix check for interactive stdout in _style function ( #47283 ) by @smart8986 in [ #47283 ] Fix get_json_schema crash on non-string docstring choices ( #47072 ) by @Sreekant13 in [ #47072 ] Make MODEL_IDS_TO_TOKENIZERS_BACKEND capture all DeepSeek R1 distills ( #47296 ) by @hmellor in [ #47296 ] Update doc preprocessing regex to prevent ReDoS ( #47187 ) by @WilliamRoyNelson in [ #47187 ] Shard on read Dtensor aware ( #46717 ) by @3outeille in [ #46717 ] Switch AMD daily CI to mi300 runners ( #47259 ) by @Abdennacer-Badaoui in [ #47259 ] tests: reduce processor test memory usage by using tiny Hub checkpoints ( #47213 ) by @ydshieh in [ #47213 ] Torch compile backend defaults to "neuron" ( #47035 ) by @michaelbenayoun in [ #47035 ] Fix flash-attn Docker build broken by setuptools 83 removing pkg_resources ( #47251 ) by @ydshieh in [ #47251 ] Add heterogeneous config support (per-layer configuration) ( #45333 ) by @eladsegal in [ #45333 ] [fix] update integration test values ( #47146 ) by @eustlb in [ #47146 ] Fix DeepSpeed SP loss aggregation and LocalLayerRepository kwargs ( #47073 ) by @sshivampeta in [ #47073 ] tests only for the top 10 download models ( #47244 ) by @3outeille in [ #47244 ] Fix InputTokensDetails missing cache_write_tokens for openai>=2.34.0 ( #47248 ) by @ydshieh in [ #47248 ] Revert "Trigger a scheduled run" ( #47249 ) by @ydshieh in [ #47249 ] Remove Executorch from CI until latest version is supported and fully tested on CI env ( #47242 ) by @IlyasMoutawwakil in [ #47242 ] Be more defensive with remap_legacy_layer_types for custom models ( #47245 ) by @hmellor in [ #47245 ] Fix DistributedConfig docstring for unimplemented sp_plan ( #47237 ) by @3outeille in [ #47237 ] Switch mlinter to 0.1.2 ( #47172 ) by @tarekziade in [ #47172 ] Trigger a scheduled run ( #47209 ) by @ydshieh in [ #47209 ] Make executorch exporter tests always use xnnpack backend ( #47201 ) by @tarekziade in [ #47201 ] No agent PR descriptions ( #45790 ) by @Rocketknight1 in [ #45790 ] Clarify that max_steps is required for datasets without len ( #47155 ) by @albertvillanova in [ #47155 ] Cleanup pipelines, stop materializing generators ( #47142 ) by @Rocketknight1 in [ #47142 ] Fix device_map computation when the no_split_modules have different sizes ( #47203 ) by @Cyrilvallez in [ #47203 ] Add native FSDP2 module + migration ( #46707 ) by @3outeille in [ #46707 ] Fix experts implementation in two spots ( #47097 ) by @remi-or in [ #47097 ] [Fix] Remove old automatic cross attn pattern from output recorders ( #47117 ) by @remi-or in [ #47117 ] 🌐 [i18n-KO] Translate accelerator_selection.md to Korean ( #47157 ) by @kkwjk2718 in [ #47157 ] [i18n-KO] Translate optimum.md to Korean and fix Furiosa typo ( #47156 ) by @kkwjk2718 in [ #47156 ] [docs] fix curly quotes rendering to straight quotes ( #47135 ) by @clijo in [ #47135 ] Fix custom code which doesn't know about the new linear layer type names ( #47174 ) by @hmellor in [ #47174 ] Reject path traversal in the transformers_weights config field ( #46890 ) by @LinZiyuu in [ #46890 ] [docs] Custom code conversion mapping ( #47114 ) by @stevhliu in [ #47114 ] Add exporters min version requirements and test skip ( #47161 ) by @IlyasMoutawwakil in [ #47161 ] tests: reduce processor test memory usage and use tiny test assets ( #47168 ) by @ydshieh in [ #47168 ] Clarify input device placement in the Quicktour inference example ( #47136 ) by @samyuktahegde in [ #47136 ] Extend continuous batching memory prediction test to XPU ( #47159 ) by @sywangyi in [ #47159 ] Fix case where _LazyAutoMapping.register is passed a str key ( #47148 ) by @hmellor in [ #47148 ] [docs] MoE decode switching ( #47149 ) by @stevhliu in [ #47149 ] add XPU output expectations for minicpm3 tests ( #47092 ) by @kaixuanliu in [ #47092 ] Diffusion gemma: fix failed test cases ( #47025 ) by @kaixuanliu in [ #47025 ] add XPU Expectation for cosmos3_omni tests ( #46880 ) by @kaixuanliu in [ #46880 ] Fix IndexError Bug in XLMRoberta/Camembert ForMultipleChoice by restoring the pooler ( #47147 ) by @pariidanDKE in [ #47147 ] Skip caching_allocator_warmup on Neuron (no reuse pool to warm; currently OOMs) ( #47029 ) by @dacorvo in [ #47029 ] [docs] continuous batching (offloading behavior, max batch tokens, block size minimum) ( #46925 ) by @stevhliu in [ #46925 ] [docs] fix autolinks ( #46968 ) by @stevhliu in [ #46968 ] revert #47121 ( #47144 ) by @eustlb in [ #47144 ] Fix output labels for AudioFlamingo3 (and related) models ( #47112 ) by @ebezzam in [ #47112 ] Fix false len claims in Trainer docstrings ( #47131 ) by @albertvillanova in [ #47131 ] processor tests: use tiny Hub repos to reduce CI memory ( #47115 ) by @ydshieh in [ #47115 ] [serge] Fix 12 integration tests for model dac failing with output_mismatch (tensor values differ (6), other (6)) ( #47121 ) by @sergereview[bot] in [ #47121 ] Fix CLI compatibility with huggingface_hub 1.22 ( #47059 ) ( #47064 ) by @dhruv7477 in [ #47064 ] we want to run the CI in the release branches ( #47125 ) by @tarekziade in [ #47125 ] Small improvement ( #47128 ) by @Cyrilvallez in [ #47128 ] [Model] Support use_cache=False for DeepSeek V4 ( #46965 ) by @kylesayrs in [ #46965 ] docs-fix: IMDb dataset link in sequence classification guide ( #47062 ) by @abhishekkapoorx in [ #47062 ] Fix AltCLIP text embedding resize test ( #47079 ) by @IMvision12 in [ #47079 ] fix mask return-type contract regression and add correctness guard for ( #47019 ) by @kaixuanliu in [ #47019 ] Fix save_pretrained with offloading and weight conversions ( #47018 ) by @Cyrilvallez in [ #47018 ] Update dev ( #47044 ) by @vasqu in [ #47044 ] [ Gemma4 ] Update 1 integration test ( #47042 ) by @vasqu in [ #47042 ] Significant community contributions The following contributors have made significant changes to the library over the last release: @ArthurZucker v5.14.0 @tarekziade ci: cover xet as well (runtime error) ( #47338 ) Pin kernels to compatible version in benchmark workflow ( #47339 ) Switch mlinter to 0.1.2 ( #47172 ) Make executorch exporter tests always use xnnpack backend ( #47201 ) Remove executorch from all-latest-gpu image + add torch smoke test ( #47196 ) we want to run the CI in the release branches ( #47125 ) @remi-or [Nit] Add kernels_fallback_ok kwarg to is_flash_attn_N_available ( #47318 ) [Nit] Add expectations for gemma4 tests on H100 ( #47311 ) [Fix] Remove deprecated argument from kernels call ( #47100 ) [Fix] Make DeepGEMM triton fallback more robust ( #47126 ) Fix experts implementation in two spots ( #47097 ) [Fix] Remove old automatic cross attn pattern from output recorders ( #47117 ) @ydshieh tests: reduce processor test memory usage by using tiny Hub checkpoints ( #47213 ) Fix flash-attn Docker build broken by setuptools 83 removing pkg_resources ( #47251 ) Fix InputTokensDetails missing cache_write_tokens for openai>=2.34.0 ( #47248 ) Revert "Trigger a scheduled run" ( #47249 ) Fix CI read-only cache failures by patching cached_files in conftest ( #47043 ) Trigger a scheduled run ( #47209 ) tests: reduce processor test memory usage and use tiny test assets ( #47168 ) processor tests: use tiny Hub repos to reduce CI memory ( #47115 ) @eladsegal Add heterogeneous config support (per-layer configuration) ( #45333 ) @eustlb [fix] update integration test values ( #47146 ) revert #47121 ( #47144 ) @Ternura143 Add TIPSv2 ( #46347 )

Read more →

Patch release v5.13.1

Patch release v5.13.1 This patch is focused on enabling transformers for the latest release of vllm! Be more defensive with remap_legacy_layer_types for custom models ( #47245 ) from @hmellor Fix custom code which doesn't know about the new linear layer type names ( #47174 ) from @hmellor Fix case where _LazyAutoMapping.register is passed a str key ( #47148 ) from @hmellor

Read more →

Release v5.13.0

Release v5.13.0 New Model additions KimiK 2.5, 2.6, and 2.7 This release includes the architecture for Kimi 2.5 which is used by 2.5-2.7: Kimi K2.5 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration. The model was proposed in Kimi K2.5: Visual Agentic Intelligence and further improved in [Kimi K2.6: Advancing Open-Source Coding](Kimi K2.5: Visual Agentic Intelligence). Kimi K2.5 achieves significant improvements on complex, end-to-end coding tasks, generalizing robustly across programming languages (Rust, Go, Python) and domains spanning front-end, DevOps, and performance optimization. The model is capable of transforming simple prompts and visual inputs into production-ready interfaces and lightweight full-stack workflows, generating structured layouts, interactive elements, and rich animations with deliberate aesthetic precision. Links: Documentation Add new model: Kimi2-6 ( #45630 ) by @zucchini-nlp in #45630 MiMo-V2-Flash MiMo-V2-Flash is a Mixture-of-Experts (MoE) language model developed by the Xiaomi MiMo team. Designed to establish a new balance between long-context modeling capabilities and inference efficiency, the model is built for strong performance in complex reasoning and agentic tasks. Trained on 27T tokens with native 32k sequence lengths, MiMo-V2-Flash seamlessly supports an extended 256K context window while significantly reducing KV-cache storage compared to standard global attention models. Links: Documentation Add Xiaomi MiMo-V2 ( #45144 ) by @casinca in #45144 Nemotron 3.5 ASR Nemotron 3.5 ASR is a 600M-parameter multilingual speech recognition model from NVIDIA, built for high-quality transcription in both low-latency streaming and high-throughput batch settings, with native punctuation and capitalization. For streaming, it offers configurable chunk sizes—80ms, 160ms, 560ms, and 1120ms, letting users trade off latency against accuracy to suit their application. Its cache-aware FastConformer-RNNT architecture is central to this capability: unlike traditional buffered streaming, which repeatedly reprocesses overlapping audio windows, the model processes only each new incoming chunk while reusing cached encoder context from prior chunks. This eliminates redundant computation, significantly improves efficiency, and minimizes end-to-end delay without sacrificing accuracy, making it well suited to real-time transcription workloads. Links: Documentation Add Nemotron 3.5 ASR Streaming ( #46565 ) by @eustlb in #46565 NemotronAsrStreaming Nemotron ASR Streaming is a 600M-parameter English speech recognition model from NVIDIA, built for high-quality transcription in both low-latency streaming and high-throughput batch settings, with native punctuation and capitalization. For streaming, it offers configurable chunk sizes—80ms, 160ms, 560ms, and 1120ms, letting users trade off latency against accuracy to suit their application. Its cache-aware FastConformer-RNNT architecture is central to this capability: unlike traditional buffered streaming, which repeatedly reprocesses overlapping audio windows, the model processes only each new incoming chunk while reusing cached encoder context from prior chunks. This eliminates redundant computation, significantly improves efficiency, and minimizes end-to-end delay without sacrificing accuracy, making it well suited to real-time transcription workloads. Links: Documentation Add Nemotron ASR Streaming ( #46332 ) by @eustlb in #46332 Qwen3 ASR Qwen3 ASR is an automatic speech recognition model from Alibaba's Qwen team that combines a Whisper-style audio encoder with a Qwen3 language model decoder for speech-to-text transcription. The model supports automatic language detection and multilingual transcription. A forced aligner model is also included. It can be used to timestamp a provided transcript and its audio. It uses the same audio encoder model with a classification head that predicts a word's length. This model can be used with the transcript from any ASR model (see the example below with Parakeet CTC). Links: Documentation Qwen3 ASR and Forced Aligner ( #43838 ) by @mbtariq82 in #43838 ZAYA ZAYA1 is a 760M active / 8.4B total parameter MoE language model trained by Zyphra. It combines Compressed Convolutional Attention (CCA), a nonlinear ZAYA1 router, and residual scaling. Links: Documentation [new model] Add Zyphra/ZAYA1-8B ( #45862 ) by @JJJYmmm in #45862 VideoPrism The VideoPrism model was proposed in the paper VideoPrism: A Foundational Visual Encoder for Video Understanding by Google DeepMind ( blog post ). VideoPrism is a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. The model is pretrained on a large-scale heterogeneous corpus containing 36M high-quality video-caption pairs and 582M video clips with noisy parallel text (e.g., ASR transcripts). The pretraining approach improves upon masked autoencoding through global-local distillation of semantic video embeddings and a token shuffling scheme, enabling the model to focus primarily on the video modality while leveraging text associated with videos. VideoPrism achieves state-of-the-art performance on 31 out of 33 video understanding benchmarks across four broad task groups, from web video question answering to computer vision for science. Links: Documentation Add Videoprism ( #39895 ) by @MHRDYN7 in #39895 RADIO RADIO (Reduce All Domains Into One) is a family of vision foundation models from NVIDIA trained by multi-teacher distillation (e.g. CLIP, DINOv2, SAM) into a single ViT backbone. It produces both an image-level summary embedding and dense spatial features , and supports variable input resolutions through a Cropped Position Embedding (CPE) patch generator. Links: Documentation Add support for RADIO models ( #46425 ) by @meatybobby in #46425 MiniCPM3 MiniCPM3 is the third-generation MiniCPM dense language model from OpenBMB. The 4B variant ( openbmb/MiniCPM3-4B ) outperforms many 7B–9B open models on standard benchmarks while remaining lightweight enough for on-device usage. MiniCPM3 combines several architectural ideas: Multi-head Latent Attention (MLA) from DeepSeek-V2, which compresses the key/value cache into a low-rank latent representation while still using rotary embeddings on a portion of the query/key heads. A standard SwiGLU MLP (no MoE). Three scalar scaling factors that govern signal flow: scale_emb — scales input embeddings. scale_depth / sqrt(num_hidden_layers) — scales residual connections. hidden_size / dim_model_base — scales hidden states before the language model head. Links: Documentation Add MiniCPM3 ( #41116 ) by @bzantium in #41116 Breaking changes A broad set of modeling changes have been made to standardize layer declarations, mask/cache construction, and hybrid-attention handling, making many models cleanly exportable (ONNX, torch.export , ExecuTorch) and fullgraph-compilable — users relying on internal modeling APIs may need to update their code accordingly. 🚨 Modeling changes for export, compile, and hybrid-attention standardization ( #46738 ) by @IlyasMoutawwakil Attention masking for image tokens in Gemma 3/4 models has been fixed to correctly respect sliding window boundaries in local layers, which changes model behavior and may affect reproducibility of previous results. 🚨 [gemma 3/4] Fix bidirectional attention masking crossing sliding window boundaries ( #46850 ) by @douglas-reid The Expert Parallelism (EP) router contract has been corrected across many models and FP8 scale format handling has been fixed, requiring users of EP or FP8 quantization with affected models to verify their configurations and potentially update conversion mappings. 🚨 EP: fix EP router contract for many models + honor FP8 scale format ( #46818 ) by @IlyasMoutawwakil The Kernels integration has been synced to the latest version, which includes a breaking change where model-type repositories are no longer accepted by the kernels interface — users must migrate to the updated kernel repository format as shown in the updated tests. 🚨 [ Kernels ] Sync to latest version ( #46039 ) by @vasqu HfExporters: Native, Unified export for PyTorch / ONNX / ExecuTorch A native, in-Transformers export pipeline — one base class ( HfExporter ), three subclasses for the runtimes we care about, one unified API: Exporter Output Runtime DynamoExporter ExportedProgram Any PyTorch runtime, AOT compilation OnnxExporter ONNXProgram Any ONNX runtime (ORT, TensorRT, OpenVINO, …) ExecutorchExporter ExecutorchProgramManager Mobile and edge (ExecuTorch) Same call shape across all three. Dynamic shapes by default. Generation-style models split automatically into prefill + decode (+ vision/audio sub-encoders for VLMs). from transformers import AutoModelForMaskedLM , AutoTokenizer from transformers . exporters import OnnxExporter , OnnxConfig model_id = "hf-internal-testing/tiny-random-BertForMaskedLM" tokenizer = AutoTokenizer . from_pretrained ( model_id ) model = AutoModelForMaskedLM . from_pretrained ( model_id ). eval () inputs = tokenizer ([ "Hello, my dog is cute" ] * 2 , return_tensors = "pt" ) onnx_program = OnnxExporter (). export ( model , inputs , config = OnnxConfig ( dynamic = True )) new_input = tokenizer ( "Hello, my cat is so adorable!" , return_tensors = "pt" ) torch . testing . assert_close ( onnx_program . call_reference ( ** new_input )[ 0 ], # numpy reference onnx_program ( ** new_input )[ 0 ], # onnxruntime rtol = 1e-4 , atol = 1e-4 , ) Swap one line for another runtime — DynamoExporter() / DynamoConfig or ExecutorchExporter() / ExecutorchConfig(backend=...) . For generative models the prefill/decode split is captured automatically: from transformers import AutoModelForCausalLM , AutoTokenizer from transformers . exporters import OnnxExporter , OnnxConfig model_id = "hf-internal-testing/tiny-random-LlamaForCausalLM" tokenizer = AutoTokenizer . from_pretrained ( model_id ) model = AutoModelForCausalLM . from_pretrained ( model_id ). eval () inputs = tokenizer ([ "Hello, my dog is cute" ] * 2 , return_tensors = "pt" ) artifacts = OnnxExporter (). export_for_generation ( model , inputs , config = OnnxConfig ( dynamic = True )) # {"prefill": ONNXProgram, "decode": ONNXProgram} # For VLMs: also vision_encoder, audio_encoder, multi_modal_projector, language_model, lm_head Kernels Kernels: Fixed a silent SDPA math-kernel fallback for GQA models with head_dim > 256 (e.g., Gemma4) that caused O(S²) memory materialization, and resolved a regression where use_kernels=True failed to apply kernel mappings. Additional improvements include lazy loading of the default kernel mapping to prevent import failures with incompatible kernel versions, ROCm routing to AITER Triton kernels for AMD GPUs, GB10/SM121 Hub-kernel support for Qwen3.6 Gated DeltaNet, and expanded documentation for the kernel API. Fix silent SDPA math-kernel fallback for GQA when key/value head_dim > 256 or differ ( #46960 ) by @Butterfingrz in [ #46960 ] [docs] AITER kernels ( #46871 ) by @stevhliu in [ #46871 ] Documentation for the kernel API ( #46754 ) by @michaelbenayoun in [ #46754 ] update kernels-community/aiter-rope version ( #46810 ) by @Abdennacer-Badaoui in [ #46810 ] Add GB10/SM121 Hub-kernel path for Qwen3.6 Gated DeltaNet ( #46423 ) by @AzeezIsh in [ #46423 ] [ Kernels ] Trigger proper kernelization on use_kernels=True ( #46755 ) by @vasqu in [ #46755 ] Lazily build the default kernel mapping to decouple kernels from normal transformers usage ( #46681 ) by @jiqing-feng in [ #46681 ] Add some AITER kernel routing for ROCm ( #46268 ) by @Abdennacer-Badaoui in [ #46268 ] fix: position ids does not exist in upstream rotary kernel ( #46619 ) by @NanoCode012 in [ #46619 ] docs(zh): add Chinese translation of kernels.md ( #46621 ) by @shoushinya123 in [ #46621 ] Generation Several generation bugs were fixed, including Mamba2 chunked-prefill and speculative decoding for hybrid models (Zamba2, Nemotron-H, Bamba, FalconH1, GraniteMoeHybrid), beam search for Mamba models, prompt lookup decoding crashes with no EOS token, and incorrect stateful model handling for LFM2. Additional improvements include reduced unnecessary generation warnings, a fix for continuous batching output mutation, and a new option to keep input tensors on CPU during generation to avoid retracing on Neuron/TPU devices. Fix Mamba2 chunked-prefill / speculative decoding for Zamba2, Nemotron-H, Bamba, FalconH1 and GraniteMoeHybrid ( #46741 ) by @Sunt-ing in [ #46741 ] Remove some unnecessary generate warnings ( #46955 ) by @Cyrilvallez in [ #46955 ] Reject assisted generation for LFM2 and LFM2-MoE (set _is_stateful) ( #46937 ) by @Sunt-ing in [ #46937 ] Fix beam search for mamba models ( #46819 ) by @Cyrilvallez in [ #46819 ] Fix prompt lookup decoding crash when no EOS token is configured ( #46790 ) by @Sunt-ing in [ #46790 ] [Continuous Batching] Snapshot generation outputs without mutating request state ( #46670 ) by @Incheonkirin in [ #46670 ] [docs] keep generation tensors on cpu ( #46675 ) by @stevhliu in [ #46675 ] feat(generation): allow user to keep input tensors on cpu ( #46590 ) by @dacorvo in [ #46590 ] Attention Several attention-related bugs were fixed in this release, including silent SDPA math-kernel fallbacks for GQA with large head dimensions, broken Flash Attention with StaticCache , incorrect causal masking in Xcodec2, a cross-attention reshape regression in Blip2, and eager GQA support in Evolla. Accelerate hook handling was also corrected for models using linear attention to prevent silently wrong results during offloading. Fix accelerate hooks for all models using linear attention ( #46978 ) by @Cyrilvallez in [ #46978 ] Fix Xcodec2 attention to be non-causal. ( #46963 ) by @ebezzam in [ #46963 ] Fix flash attention with StaticCache ( #46914 ) by @Cyrilvallez in [ #46914 ] Fix Evolla eager attention for the GQA text decoder ( #46860 ) by @jiqing-feng in [ #46860 ] [docs] metal flash attention ( #46349 ) by @stevhliu in [ #46349 ] [ Blip2 ] Fix cross attention reshape ( #46695 ) by @vasqu in [ #46695 ] Cache Cache APIs were improved by consolidating redundant getters into a cleaner get_max_length method and updating documentation accordingly. Several bug fixes were also applied, including correcting mask generation beyond sliding windows, fixing a dimension issue in cumulative length tracking, resolving device mismatches in offloaded cache for hybrid models, and fixing crashes when loading trust_remote_code models from symlinked local caches. [docs] update cache apis ( #46892 ) by @stevhliu in [ #46892 ] Rework some old cache getters/properties ( #46862 ) by @Cyrilvallez in [ #46862 ] Fix expanded dim in the cache's cumulative length ( #46856 ) by @Cyrilvallez in [ #46856 ] Fix mask when generating beyond sliding window ( #46839 ) by @zucchini-nlp in [ #46839 ] Fix offloaded cache device mismatch on hybrid models ( #46748 ) by @Sunt-ing in [ #46748 ] Fix dynamic module symlinked cache on trust_remote_code models ( #46618 ) by @ldkhang1201 in [ #46618 ] Serve Several fixes and improvements were made to the Serve functionality, including lazy imports to prevent CLI crashes when the optional serve extra is not installed, a fix for dropped attributes during serialization of subclassed Pydantic models, and added documentation for the kernel API. fix(cli/serve): import serve handlers lazily so the CLI works without the serve extra ( #46473 ) by @ in [ #46473 ] [Fix] Serve drops some attributes at serialization ( #46680 ) by @remi-or in [ #46680 ] Reduce per_page from 100 to 50 in GitHub API calls to avoid server errors ( #46678 ) by @ydshieh in [ #46678 ] Quantization Fixed dtype casting bugs in Gemma4's vision and audio multimodal embedders when using BitsAndBytes quantization, where inputs were incorrectly cast to integer storage dtypes ( uint8 / int8 ) instead of the actual compute dtype. Also corrected FP8 quantization to round block scales before quantizing weights, ensuring dequantization produces correct values for ue8m0 (DeepSeek-V4 style) format. [Gemma4] Fix dtype casting for quantized vision/audio embedders ( #46933 ) by @sharmax-vikas in [ #46933 ] Fix dtype casting for quantized multimodal embedders ( #46904 ) by @praful-srinivasan-027 in [ #46904 ] Round the ue8m0 FP8 scale before quantizing so dequant matches the stored inverse ( #46763 ) by @Incheonkirin in [ #46763 ] Bugfixes and improvements Update workflow callers to use transformers-ci ( #47040 ) by @ydshieh in [ #47040 ] Add HunYuan VL model ( #46417 ) by @Mi-Jiazhi in [ #46417 ] Add tiny_model_id support to ProcessorTesterMixin for memory-sensitive tests ( #47005 ) by @ydshieh in [ #47005 ] chore(linter): add TRF018 modeling rule ( #46259 ) by @tarekziade in [ #46259 ] [PoC] HF exporters ( #41992 ) by @IlyasMoutawwakil in [ #41992 ] TST Skip PEFT tests if PEFT version is too low ( #47027 ) by @BenjaminBossan in [ #47027 ] CI Add PEFT integration tests ( #47021 ) by @BenjaminBossan in [ #47021 ] [glm-mode-dsa] Indexer uses interleaved rope ( #46842 ) by @pcuenca in [ #46842 ] Use standard arg names in Mllama ( #46977 ) by @zucchini-nlp in [ #46977 ] Bump min peft 0.19.1 remove weight conversion duplicate code ( #46442 ) by @BenjaminBossan in [ #46442 ] Raise a loud error for missing prefix ( #46980 ) by @Rocketknight1 in [ #46980 ] Fix typo in Qwen3 ASR no_split_module ( #47002 ) by @ebezzam in [ #47002 ] only in the original repo ( #46982 ) by @tarekziade in [ #46982 ] Fix typos in Gemma 4 Assistant documentation ( #46975 ) by @RaunaqDavidNath in [ #46975 ] the CI status should be a comment ( #46976 ) by @tarekziade in [ #46976 ] QwenVL model conversion ( #46881 ) by @zucchini-nlp in [ #46881 ] Remove default dtype in FusedRMSNormGated modules ( #46953 ) by @Cyrilvallez in [ #46953 ] FIX PEFT test changed error type ( #46959 ) by @BenjaminBossan in [ #46959 ] Fix path traversal via vocab-file arguments in tokenizer_config.json ( #46279 ) by @LinZiyuu in [ #46279 ] docs(conditional_detr): fix num_queries default in docstring (100 -> 300) ( #46939 ) by @Kropiunig in [ #46939 ] Use common floats_list method for feature extractor tests. ( #46956 ) by @ebezzam in [ #46956 ] Fix RT-DETR indexing error when num_feature_levels exceeds backbone o… ( #46833 ) by @c1prk in [ #46833 ] Fix Florence2 training-loss double-shift (same pattern as Moonshine #… ( #46898 ) by @sharmax-vikas in [ #46898 ] [Olmo3] different RoPE per layer type ( #46911 ) by @zucchini-nlp in [ #46911 ] Use inspect.getsource instead of open() for source-reading in can_set *_implementation ( #46207 ) by @rasmi in [ #46207 ] Don't pin the gated delta net norm to cuda:0 with a hardcoded device ( #46817 ) by @Sunt-ing in [ #46817 ] Fix auto-mappings registration for remote code & fixes a few custom code issues ( #46876 ) by @Cyrilvallez in [ #46876 ] Fix broken internal documentation links ( #46945 ) by @sezer-muhammed in [ #46945 ] Insert a Grafana badge in the PR ( #46774 ) by @tarekziade in [ #46774 ] [NemotronAsrStreaming] fix pipeline ( #46870 ) by @eustlb in [ #46870 ] [NemotronAsrStreaming] processor without modular ( #46865 ) by @eustlb in [ #46865 ] [ Dia ] Fix docs ( #46923 ) by @vasqu in [ #46923 ] [Docs] Fix full disk offloading docs ( #46905 ) by @kylesayrs in [ #46905 ] [CB] Changes to increase max_batch_tokens ( #46712 ) by @remi-or in [ #46712 ] Redirect to diffusers pipe in docs for experimental features ( #46875 ) by @zucchini-nlp in [ #46875 ] Install in docker ( #46910 ) by @ydshieh in [ #46910 ] [CI] Use pre-computed _OLD_MODELS in test_new_models_require_torchvision_backend ( #46882 ) by @ydshieh in [ #46882 ] call transformers-ci in a nightly run ( #46811 ) by @tarekziade in [ #46811 ] [docs] full disk offloading ( #46893 ) by @stevhliu in [ #46893 ] TST Run fast PEFT tests in normal CI ( #45679 ) by @BenjaminBossan in [ #45679 ] nemotron_asr_streaming: set _supports_flex_attn to False ( #46878 ) by @kaixuanliu in [ #46878 ] Add native masked MSE loss for Sapiens2ForPoseEstimation ( #46764 ) by @Sainava in [ #46764 ] blip 2 fix ( #46816 ) by @itazap in [ #46816 ] Use meshgrid for brevity ( #46861 ) by @zucchini-nlp in [ #46861 ] Add xcodec2 model ( #44178 ) by @ebezzam in [ #44178 ] Prevent auto-class from being modified for all models ( #46844 ) by @zucchini-nlp in [ #46844 ] Add Spanish translation of the torch.compile page ( #46852 ) by @delcenjo in [ #46852 ] docs: Update NeMo AutoModel doc examples ( #46857 ) by @adil-a in [ #46857 ] [docs] distributed training ( #44420 ) by @stevhliu in [ #44420 ] [docs] require trust_remote_code for custom_generate ( #46677 ) by @stevhliu in [ #46677 ] add distributed config ( #46705 ) by @3outeille in [ #46705 ] [Offloading] [Bugfix] Fix disk offloading of models with explicit tensor dtypes ( #46849 ) by @kylesayrs in [ #46849 ] Streamable chat parsing ( #45847 ) by @Rocketknight1 in [ #45847 ] Fix BitNet packed-weight unpacking dtype ( F.linear dtype mismatch) ( #46808 ) by @jiqing-feng in [ #46808 ] Fix typos in code ( #46579 ) by @cyyever in [ #46579 ] Fix Moonshine training-loss double-shift (train against labels, not labels[..., 1:]) ( #46784 ) by @Incheonkirin in [ #46784 ] [CB] Fix issues with FA read / writes ( #46765 ) by @remi-or in [ #46765 ] Switch decorator order ( #46853 ) by @Cyrilvallez in [ #46853 ] docs(trainer): add JIT checkpointing to trainer recipes ( #46826 ) by @efazal in [ #46826 ] Import diffusion_gemma in models init ( #46841 ) by @boringcrypto in [ #46841 ] [skills] help your agent get started ( #45732 ) by @stevhliu in [ #45732 ] Fix use_cache with seq_len > 1 ( #46032 ) ( #46084 ) by @Ramshankar07 in [ #46084 ] [Offloading] Support full disk offloading ( #46749 ) by @kylesayrs in [ #46749 ] fix: raise ValueError for empty conversation in apply_chat_template ( #46753 ) by @sharmax-vikas in [ #46753 ] Fix VideoPrismForVideoClassification returning last_hidden_state as h… ( #46830 ) by @sharmax-vikas in [ #46830 ] Avoid NumPy 2.0 __array__ copy-keyword deprecation in create_mm_token_type_ids ( #46827 ) by @qgallouedec in [ #46827 ] docs: update apple silicon doc with safetensors 0.8.0 benefits ( #46744 ) by @McPatate in [ #46744 ] [ CB ] Add FA2 to the fast path ( #46729 ) by @vasqu in [ #46729 ] Fix flex_attention block mask creation when get_seq_length returns a tensor ( #46802 ) by @jiqing-feng in [ #46802 ] Fix left-padding token selection in BioGptForSequenceClassification ( #46782 ) by @Sunt-ing in [ #46782 ] Fix broken internal links in model documentation ( #46807 ) by @ShamSaleem in [ #46807 ] DiffusionGemma: mask layout and CI ( #46654 ) by @zucchini-nlp in [ #46654 ] Use cached added-token dicts in per-token decode loops ( #46535 ) by @ishan-1010 in [ #46535 ] fix another flaky test ( #46767 ) by @zucchini-nlp in [ #46767 ] Fix secondary rate limit when downloading artifacts in slack report ( #46796 ) by @ydshieh in [ #46796 ] docs: move SmolLM3 to Text models category in _toctree.yml ( #46770 ) by @yyouretoast in [ #46770 ] Fix several bugs in cache_implementation=static ( #46446 ) by @dacorvo in [ #46446 ] [CI] Fix artifact download path in self-comment-ci workflow ( #46769 ) by @ydshieh in [ #46769 ] fixes per head minimaxm3 ( #46719 ) by @ArthurZucker in [ #46719 ] [ CI ] Fix some failures introduced by myself 😬 ( #46751 ) by @vasqu in [ #46751 ] Fix regression in ProcessorMixin._load_tokenizer_from_pretrained for tokenizers at root ( #46592 ) by @ in [ #46592 ] fix(aria): use math.ceil in get_number_of_image_patches to match actual patch count ( #46732 ) by @arnavkewalram in [ #46732 ] Return logits from semantic segmentation post-process ( #46163 ) by @guarin in [ #46163 ] Fall back to the for-loop grouped_mm on CPU ( #46743 ) by @Sunt-ing in [ #46743 ] Kernelize refactor ( #46520 ) by @michaelbenayoun in [ #46520 ] ci: add comment explaining why secrets are not inherited in security gate ( #46750 ) by @ydshieh in [ #46750 ] ci: trigger PR CI on ci-* branches ( #46746 ) by @ydshieh in [ #46746 ] finegrained v3 ( #46742 ) by @IlyasMoutawwakil in [ #46742 ] Improve AutoImageProcessor error for unavailable backends ( #46727 ) by @sisaman in [ #46727 ] skip decorators must appear after @parameterized.expand in pytest ( #46737 ) by @rasmi in [ #46737 ] [RecurrentGemma] Support attn_implementation dispatch ( #46320 ) by @YangKai0616 in [ #46320 ] [docs] clarify initialization module usage ( #46698 ) by @stevhliu in [ #46698 ] feat: bump safetensors to 0.8.0 ( #46523 ) by @McPatate in [ #46523 ] ci: disable CircleCI by replacing config with no-op ( #46721 ) by @ydshieh in [ #46721 ] [CB] Fix offloading ( #46587 ) by @remi-or in [ #46587 ] [ Templates ] Update members ( #46720 ) by @vasqu in [ #46720 ] feat[vLLM x v5]: Expose max_source_positions on VibeVoiceAsrConfig ( #46472 ) by @harshaljanjani in [ #46472 ] Laguna: support per-element output gating ( #46690 ) by @joerowell in [ #46690 ] ci: grant pull-requests:write to the security gate caller ( #46715 ) by @ydshieh in [ #46715 ] Multi-gpu loading when the whole backbone is tied ( #46625 ) by @zucchini-nlp in [ #46625 ] Delete docstring if same as in auto-doc ( #46284 ) by @zucchini-nlp in [ #46284 ] Update GLM-5.2 docs ( #46703 ) by @Dovis01 in [ #46703 ] add conversion scripts for EUPE ( #46691 ) by @molbap in [ #46691 ] [docs] compile level and batch/scheduling limits ( #46676 ) by @stevhliu in [ #46676 ] [blip_2] Support attn_implementation dispatch ( #46401 ) by @YangKai0616 in [ #46401 ] [CTRL] Support attn_implementation dispatch ( #46073 ) by @YangKai0616 in [ #46073 ] Lfm2: also thread seq_idx through ShortConv.slow_forward (non-fast-path) ( #46633 ) by @ChangyiYang in [ #46633 ] feat(pipelines): accept numpy arrays and tensors in ImageClassificationPipeline ( #39607 ) ( #46573 ) by @kamran-nizamani in [ #46573 ] Smovlm: pad videos up to max frames ( #46662 ) by @zucchini-nlp in [ #46662 ] mistral common backend fix ( #46667 ) by @itazap in [ #46667 ] [pr template] update ( #46606 ) by @stevhliu in [ #46606 ] Fix AttributeError in auto_factory when model_class lacks config_class ( #46669 ) by @atharv1945 in [ #46669 ] [CB] Slice logits inside the model ( #46660 ) by @remi-or in [ #46660 ] ci: add NO_COLOR=1 to suppress ANSI color codes in CI output ( #46659 ) by @ydshieh in [ #46659 ] Fix dynamic RoPE not resetting inv_freq when layer_type is None ( #46624 ) by @Incheonkirin in [ #46624 ] Better processing tests ( #46374 ) by @zucchini-nlp in [ #46374 ] ci: add merge_group trigger to pr-ci-caller.yml ( #46668 ) by @ydshieh in [ #46668 ] skip invalid quant_cache test for nemotron_h ( #46368 ) by @kaixuanliu in [ #46368 ] Revert "Disable PR CI workflow for PRs from forked repo. during the weekend" ( #46652 ) by @ydshieh in [ #46652 ] [CB] Fix seqlens and use TypedDict ( #46593 ) by @remi-or in [ #46593 ] Disable PR CI workflow for PRs from forked repo. during the weekend ( #46609 ) by @ydshieh in [ #46609 ] Update post release ( #46608 ) by @vasqu in [ #46608 ] Fix peft lower bound ( #46605 ) by @hmellor in [ #46605 ] Fix docstring formatting issues causing Sphinx autodoc warnings ( #46596 ) by @kurtmckee in [ #46596 ] Significant community contributions The following contributors have made significant changes to the library over the last release: @ydshieh Update workflow callers to use transformers-ci ( #47040 ) Add tiny_model_id support to ProcessorTesterMixin for memory-sensitive tests ( #47005 ) Install in docker ( #46910 ) [CI] Use pre-computed _OLD_MODELS in test_new_models_require_torchvision_backend ( #46882 ) Fix secondary rate limit when downloading artifacts in slack report ( #46796 ) [CI] Fix artifact download path in self-comment-ci workflow ( #46769 ) ci: add comment explaining why secrets are not inherited in security gate ( #46750 ) ci: trigger PR CI on ci-* branches ( #46746 ) ci: disable CircleCI by replacing config with no-op ( #46721 ) ci: grant pull-requests:write to the security gate caller ( #46715 ) Reduce per_page from 100 to 50 in GitHub API calls to avoid server errors ( #46678 ) ci: add NO_COLOR=1 to suppress ANSI color codes in CI output ( #46659 ) ci: add merge_group trigger to pr-ci-caller.yml ( #46668 ) Revert "Disable PR CI workflow for PRs from forked repo. during the weekend" ( #46652 ) Disable PR CI workflow for PRs from forked repo. during the weekend ( #46609 ) @Mi-Jiazhi Add HunYuan VL model ( #46417 ) @tarekziade chore(linter): add TRF018 modeling rule ( #46259 ) only in the original repo ( #46982 ) the CI status should be a comment ( #46976 ) Insert a Grafana badge in the PR ( #46774 ) call transformers-ci in a nightly run ( #46811 ) @casinca Add Xiaomi MiMo-V2 ( #45144 ) @JJJYmmm [new model] Add Zyphra/ZAYA1-8B ( #45862 ) @ebezzam Fix typo in Qwen3 ASR no_split_module ( #47002 ) Fix Xcodec2 attention to be non-causal. ( #46963 ) Use common floats_list method for feature extractor tests. ( #46956 ) Add xcodec2 model ( #44178 ) @meatybobby Add support for RADIO models ( #46425 ) @douglas-reid 🚨 [gemma 3/4] Fix bidirectional attention masking crossing sliding window boundaries ( #46850 ) @Sunt-ing Fix Mamba2 chunked-prefill / speculative decoding for Zamba2, Nemotron-H, Bamba, FalconH1 and GraniteMoeHybrid ( #46741 ) Reject assisted generation for LFM2 and LFM2-MoE (set _is_stateful) ( #46937 ) Don't pin the gated delta net norm to cuda:0 with a hardcoded device ( #46817 ) Fix prompt lookup decoding crash when no EOS token is configured ( #46790 ) Fix left-padding token selection in BioGptForSequenceClassification ( #46782 ) Fix offloaded cache device mismatch on hybrid models ( #46748 ) Fall back to the for-loop grouped_mm on CPU ( #46743 ) @eustlb Add Nemotron 3.5 ASR Streaming ( #46565 ) [NemotronAsrStreaming] fix pipeline ( #46870 ) [NemotronAsrStreaming] processor without modular ( #46865 ) Add Nemotron ASR Streaming ( #46332 ) [fix] enable base64 str audio in load_audio ( #46694 ) @vasqu [ Dia ] Fix docs ( #46923 ) [ CB ] Add FA2 to the fast path ( #46729 ) [ Kernels ] Trigger proper kernelization on use_kernels=True ( #46755 ) [ CI ] Fix some failures introduced by myself 😬 ( #46751 ) 🚨 [ Kernels ] Sync to latest version ( #46039 ) [ Templates ] Update members ( #46720 ) [ Blip2 ] Fix cross attention reshape ( #46695 ) Update post release ( #46608 ) @mbtariq82 Qwen3 ASR and Forced Aligner ( #43838 ) @remi-or [CB] Changes to increase max_batch_tokens ( #46712 ) [CB] Fix issues with FA read / writes ( #46765 ) [CB] Fix offloading ( #46587 ) [Fix] Serve drops some attributes at serialization ( #46680 ) [CB] Slice logits inside the model ( #46660 ) [CB] Fix seqlens and use TypedDict ( #46593 ) @jiqing-feng Fix BitNet packed-weight unpacking dtype ( F.linear dtype mismatch) ( #46808 ) Fix Evolla eager attention for the GQA text decoder ( #46860 ) Fix flex_attention block mask creation when get_seq_length returns a tensor ( #46802 ) Lazily build the default kernel mapping to decouple kernels from normal transformers usage ( #46681 ) @bzantium Add MiniCPM3 ( #41116 ) @MHRDYN7 Add Videoprism ( #39895 ) @YangKai0616 [RecurrentGemma] Support attn_implementation dispatch ( #46320 ) [blip_2] Support attn_implementation dispatch ( #46401 ) [CTRL] Support attn_implementation dispatch ( #46073 )

Read more →

Patch release v5.10.4

Patch release v5.10.4 Update: Note that on pypi 5.10.3 doesn't exist and this this saved under 5.10.4 (so essentially a minor version skipped). Sorry about that, that's on me. Just wanted to clarify to make this less confusing! A few fixes needed for vLLM to sync with transformers 🤗 [fix] regression introduced by #45534 #46456 by @eustlb ( #46456 ) Fix {image/video/audio}_token_ids in ProcessorMixin #46500 by @hmellor ( #46500 ) Fix InternVL models #46524 by @hmellor ( #46524 ) Fix the offsets in processing #46525 by @zucchini-nlp ( #46525 ) Fix peft lower bound #46605 by @hmellor ( #46605 ) mistral common backend fix #46667 by @itazap ( #46667 ) Full Changelog : v5.10.2...v5.10.3

Read more →

Patch release v5.12.1

Patch release v5.12.1 Updated the lower bound for PEFT and a fix for auto tokenizer to properly resolve the mistral tokenizer (when mistral-common is installed). This is similar to v.5.10.3 minus the fixes that were already included in the main release - vLLM will first target 5.10.3 🤗 Fix peft lower bound #46605 by @hmellor ( #46605 ) mistral common backend fix #46667 by @itazap ( #46667 ) Full Changelog : v5.12.0...v5.12.1

Read more →

Uniform Changelog API

Access Hugging Face changelog updates through our uniform API. Same JSON structure across all sources — no adapter-specific parsing needed.

API Endpoint
GET https://watchchangelog.com/api/v1/entries?source=huggingface.releases
Response Sample
{
  "source": "huggingface.releases",
  "vendor": "Hugging Face",
  "id": "tag:github.com,2008:Repository/155220641/v5.16.0",
  "published_at": "2026-08-26T15:03:33.000Z",
  "title": "Release: v5.16.0",
  "url": "https://github.com/huggingface/transformers/releases/tag/v5.16.0",
  "summary": "Release v5.16.0 New Model additions Qwen4-Exp Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE). GR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm. It mixes multiple residual streams with fine-grained elementwise gating before each attention and Mixture-of-Experts (MoE) block, then controls how much of the block output is injected back into each stream. QSA uses multiple query heads to score compressed key blocks, selects the most relevant contiguous token blocks, and keeps the incomplete trailing block uncompressed. This block-level selection reduces indexing overhead and improves memory locality for long sequences. Combined with Gated DeltaNet, QSA makes Qwen4-Exp the first hybrid architecture to integrate linear and sparse attention, substantially improving inference efficiency for long-context workloads. PLE enriches selected decoder layers with layer-specific lexical features derived from hashed token n-grams and a dilated depthwise convolution. Links: Documentation Add Qwen4Exp model ( #48337 ) by @Cyrilvallez in #48337 GraniteSpeech5 Granite Speech 5.0 Turbo CTC is a lightweight (~470M parameters) conformer encoder for automatic speech recognition, trained with Connectionist Temporal Classification (CTC) on BPE targets. It is a fast, encoder-only member of the Granite Speech family: transcription requires a single forward pass followed by greedy CTC decoding, with no autoregressive decoder. Architecturally, it extends the Granite Speech conformer CTC encoder with: Frame stacking + block-wise time subsampling : the feature extractor stacks pairs of log-mel(+delta) frames (2x), and the first two conformer blocks each subsample time by 2 through a stride-2 depthwise convolution (with a mean-pooled residual), for a total 8x time reduction at 10 ms mel hop. Block attention with Shaw's relative positional embeddings : attention is computed over fixed-size blocks (the sequence is right-padded to a whole number of blocks, with padded frames masked out), using separate bias-free query/key/value projections. Self-conditioned CTC : the CTC posteriors of the middle layer are projected and fed back into the hidden states, and the CTC head is shared between this mid-layer self-conditioning and the final prediction. Links: Documentation Add Granite Speech 5.0 - ( #48288 ) by @eustlb in #48288 Step3p7 Step-3.7-Flash was proposed in Step 3.7 Flash by StepFun. It is a 198B-parameter sparse Mixture-of-Experts vision-language model, pairing a 196B-parameter MoE language backbone with a 1.8B-parameter vision encoder for native image understanding. StepFun hasn't published a technical report for Step-3.7-Flash, so the details below are drawn from the released checkpoint's configuration rather than a paper. Sparse MoE decoder : all but the first 3 decoder layers route through a MoE block of 288 routed experts (top-8 per token) plus a single shared expert. The router scores experts with a sigmoid and a learned per-expert bias instead of an auxiliary load-balancing loss, the same strategy as DeepSeek-V3 . Gated attention : each attention layer adds an extra projection whose sigmoid output gates the attention output per head, before the output projection — the same Gated Attention mechanism used in Qwen3-Next . A subset of layers use fewer heads and a sliding window instead of full attention. Multi-token prediction : some checkpoints ship extra decoder layers trained for multi-token prediction, which [ ~GenerationMixin.generate ] can use for speculative decoding via use_mtp=True . Vision encoder : a SigLIP-style ViT with 2-D rotary position embeddings and a learned per-layer scale on the attention and MLP branches. Its output is downsampled 4x by two stride-2 convolutions before a linear projector maps it into the text model's hidden size. Dynamic image tiling : instead of a fixed tile grid, the image processor picks its tiling window from each image's own aspect ratio, producing one downscaled global view plus zero or more local high-resolution crops per image. Links: Documentation [new model] step 3.7 ( #46658 ) by @itazap in #46658 CohereCompass CohereCompass is the base architecture for small, specialized (vision-)language models trained by Cohere. Links: Documentation Add CohereCompass modeling ( #47878 ) by @calpt in #47878 ESMC and ESMFold2 ESMC and ESMFold2 are new state-of-the-art protein language and folding models from BioHub. ESMC is trained with a masked language modeling objective, and it can be easily transferred to sequence and token classification tasks for proteins. Checkpoints exist in various sizes, from 300M parameters up to 6B parameters. It works as a drop-in replacement for older ESM-2 and ESM-3 models, with significantly higher accuracy. ESMFold2 is a state-of-the-art protein folding model which produces high accuracy predictions. It uses an iterated diffusion approach that is significantly different from the original ESMFold, offering huge improvements in accuracy for more complex structures. Links: Documentation ESMC , Documentation ESMFold2 Port ESMC and ESMFold2 to Transformers ( #46419 ) by @Rocketknight1 in #46419 Breaking changes The legacy tensor-parallel implementation has been replaced with a DTensor-native backend, so users relying on the previous TP API for inference or training must migrate to the new DTensor-based interface. 🚨 TP dtensor API inference + training ( #47579 ) by @3outeille attn_implementation=\"sdpa\" dispatch is now properly supported for wav2vec2_conformer, wav2vec2-bert, and SeamlessM4T/v2 models, which may change initialization behavior for users who previously worked around this limitation. 🚨[wav2vec2] Support attn_implementation=sdpa dispatch ( #46196 ) by @YangKai0616 FuyuProcessor no longer returns the image_patch_indices output, so any code that depends on this field must be updated to remove references to it. 🚨 Leftover processors ( #47924 ) by @zucchini-nlp Cache Several cache-related bugs were fixed in this release, including an off-by-one error in the sliding window cache, Whisper speculative decoding cache corruption, CpmAnt use-cache failures, Qwen2.5-Omni/Qwen3-Omni-MoE generation with compilable caches, and compressed-tensors loading for KV-cache-only quantized models. Documentation was also added for cache token removal using negative values, and per-layer cache configuration support (allowing models to use different cache settings per layer) was introduced. Cpmant fix use cache ( #48013 ) by @jiqing-feng in [ #48013 ] [docs] Cache crop ( #47950 ) by @stevhliu in [ #47950 ] Revert \"Support per-layer cache configuration and attention-mask selection\" ( #48175 ) by @Cyrilvallez in [ #48175 ] Support per-layer cache configuration and attention-mask selection ( #47901 ) by @eladsegal in [ #47901 ] Fix Qwen2.5-Omni / Qwen3-Omni-MoE generation with a compilable cache ( #47872 ) by @jiqing-feng in [ #47872 ] Fix sliding window cache index off-by-one on wraparound ( #47708 ) by @hameedibrh in [ #47708 ] [Whisper] Fix speculative decoding: UnboundLocalError, cache corruption, and speed regression ( #48000 ) by @ydshieh in [ #48000 ] Fix compressed-tensors loading for KV-cache-only quantized models ( #47904 ) by @kylesayrs in [ #47904 ] Generation This release fixes several generation bugs across multiple models, including Whisper speculative decoding issues (UnboundLocalError, cache corruption, speed regression, and left-padded batch position IDs), broken image generation in Emu3, garbage output in OLMo/GPTNeoX, and Qwen2.5-Omni/Qwen3-Omni-MoE generation with compilable caches. Additionally, logit distributions for candidate generators using sampling are now aligned by returning logits after applying logit processors. [Whisper] Fix speculative decoding: preserve cleared suppress tokens through super().generate() ( #48108 ) by @ydshieh in [ #48108 ] [Whisper] Fix decoder position IDs for left-padded batches in longform generation ( #48028 ) by @ydshieh in [ #48028 ] [serge] Fix 2 integration tests for model generation failing with import_or_config (other (2)) ( #48061 ) by @sergereview[bot] in [ #48061 ] Align logit distributions for CandidateGenerators using sampling ( #48007 ) by @Cyrilvallez in [ #48007 ] [GPTNeoX] Fix post_processor not overridden when loading from pretrained (OLMo garbage generation) ( #47988 ) by @ydshieh in [ #47988 ] [emu3] 🦮 Black Labrador is back! Fix image generation broken since #37033 ( #47948 ) by @ydshieh in [ #47948 ] Attention Several attention-related bug fixes were made in this release, including correcting a SigLIP2 documentation typo, fixing Flash/SDPA attention dispatch tests for xcodec2 and ROCm RDNA GPUs, resolving a GPT2 cross-attention mask being silently discarded, and enabling SDPA support declaration in TimmWrapper . Per-layer cache configuration and attention-mask selection support was also introduced, allowing models with heterogeneous layer configurations to use distinct sliding_window , attention_chunk_size , and number_of_conv_states values per layer. doc: Fix typo in SigLIP2 Flash Attention code example ( #48197 ) by @VimalN2005 in [ #48197 ] [xcodec2] Fix flex attention and flash dispatch tests ( #48244 ) by @jiqing-feng in [ #48244 ] Fix ROCm SDPA-flash skip guard that crashes on RDNA GPUs ( #47965 ) by @Abdennacer-Badaoui in [ #47965 ] [GPT2] Fix encoder_attention_mask being silently discarded in cross-attention ( #47946 ) by @DavidJohnQuinlan in [ #47946 ] Declare sdpa support in TimmWrapper ( #47939 ) by @jiqing-feng in [ #47939 ] Quantization Quantization improvements include adding NVFP4 quantization support via HF kernels (enabling on-the-fly BF16 weight quantization with ~50% memory reduction), and fixing several bugs: reverting a regression in is_quantization_compressed that caused incorrect module layouts for packed-format checkpoints, fixing CLIP weight initialization failures with quantized checkpoints, and restoring KV-cache quantization setup for KV-cache-only quantized models. Revert \"[Quantization]: Refactor is_quantization_compressed for format-based detection\" ( #48072 ) by @subin9 in [ #48072 ] feat: add nvfp4 quantization ( #47883 ) by @drbh in [ #47883 ] [DeepSeekV2] Fix integration tests OOM: use device_map=auto instead of 8-bit quantization ( #47991 ) by @ydshieh in [ #47991 ] Fix CLIP _init_weights when a child module carries quantized weights ( #47921 ) by @Bluear7878 in [ #47921 ] Parallelization Introduced a naive pipeline parallel inference engine supporting tied/untied weight embeddings with seamless generate() integration, while restoring backward compatibility for the tensor-parallel API with a deprecation cycle for tp_plan in from_pretrained() . Additionally fixed a model parallel bug in the BLT model affecting beam search. Restore BC for the tensor-parallel API ( #48300 ) by @ArthurZucker in [ #48300 ] fix bug for blt model parallel bug ( #48327 ) by @kaixuanliu in [ #48327 ] Pipeline parallel naive inference ( #47289 ) by @3outeille in [ #47289 ] Kernels Kernel support was improved with documentation updates highlighting supported models, a fix for export crashes on kernel-decorated functions by adding a is_torchdynamo_exporting guard, and the default Flash Attention 2 hub kernel version was bumped to v3 to resolve compatibility issues with newer PyTorch versions. [docs] Kernel supported models ( #48258 ) by @stevhliu in [ #48258 ] [Fix] Export crashes on kernel-decorated function ( #47808 ) by @remi-or in [ #47808 ] Bump default flash-attn2 hub kernel version to v3 ( #47863 ) by @jiqing-feng in [ #47863 ] Bugfixes and improvements Fix video-llama modular conversion ( #48336 ) by @zucchini-nlp in [ #48336 ] CI: gate the hunyuan-moe slow test ( #48330 ) by @tarekziade in [ #48330 ] Add a regression test for force_accelerate_hooks signature preservation ( #48260 ) by @wtdcode in [ #48260 ] Add an opt-in per-frame pixel cap (cap_pixels_per_frame) to the Qwen3-VL video processor ( #48071 ) by @dkrisman in [ #48071 ] Fix scores type in stopping criteria docstrings ( #47676 ) by @qgallouedec in [ #47676 ] Docstring check didn't match some file - fix it ( #48121 ) by @zucchini-nlp in [ #48121 ] Add shared ImageProcessingTester ( #47745 ) by @guarin in [ #47745 ] Fix AutoTokenizer returning TokenizersBackend for DeepSeek-R1-Distill-Qwen models ( #48211 ) by @ydshieh in [ #48211 ] [docs] Fix failing doctests ( #47687 ) by @stevhliu in [ #47687 ] Fix build_2d_sinusoidal_position_embedding on MPS ( #47897 ) by @guarin in [ #47897 ] [Gemma4] Investigate flaky test_generation_beyond_sliding_window_1_eager ( #48236 ) by @ydshieh in [ #48236 ] Let gradient checkpointing skip layers with every_n_layers ( #48200 ) by @qgallouedec in [ #48200 ] Disable daily nightly CI ( #48292 ) by @remi-or in [ #48292 ] gs ( #48288 ) by @eustlb in [ #48288 ] CI: fix muse OOMs ( #48284 ) by @tarekziade in [ #48284 ] Fix BayesianDetectorModel.from_pretrained() by calling post_init() ( #48254 ) by @woojinpaik in [ #48254 ] [Fix] Avoid duplicating tests in CI ( #48287 ) by @remi-or in [ #48287 ] [ GDN ] Fix recurrent FLA fallback ( #48266 ) by @vasqu in [ #48266 ] Fix tie_word_embeddings not lifted from text_config for some VLM configs (BC regression) ( #45857 ) by @qgallouedec in [ #45857 ] replace xpu-smi subprocess call in benchmark_v2 ( #48083 ) by @kaixuanliu in [ #48083 ] ignore mlinter ci file ( #48267 ) by @tarekziade in [ #48267 ] Fix dtype mismatch in grouped_mm_fallback for LoRA training on Mamba+… ( #47933 ) by @adh-aakriti in [ #47933 ] Compute MoE load-balancing loss per layer to avoid giant one-hot materialization −99.7% @ 128k ( #48131 ) by @qgallouedec in [ #48131 ] deterministic layer_types buffer registration in multiple models ( #48162 ) by @mowoe in [ #48162 ] Fix nemotron_h save_pretrained emitting singular backbone.embedding.weight ( #48075 ) by @yuekaizhang in [ #48075 ] force_accelerate_hooks should not hide the signature it wraps ( #48156 ) by @SunMarc in [ #48156 ] fix(data_collator): align TokenClassification numpy_call with torch_call ( #48212 ) by @ in [ #48212 ] Fix a typo in a use of a local variable field_ in a test ( #48184 ) by @AleksMat in [ #48184 ] [serge] Fix 2 integration tests regressed by commit 16780c8 (PR #47622 ) ( #48134 ) by @sergereview[bot] in [ #48134 ] [Gemma4] Fix stale expected values in integration tests ( #48233 ) by @ydshieh in [ #48233 ] [TableTransformer, PI0] Fix stale expected values and OOM in integration tests ( #48198 ) by @ydshieh in [ #48198 ] Post two CI badges on a PR: CPU PR CI and GPU run-slow ( #48190 ) by @tarekziade in [ #48190 ] Fix stale expected values in integration tests (cuda sm_86 / Aug04 regressions) ( #48171 ) by @ydshieh in [ #48171 ] Assign a reviewer even when a codeowner has left, and route models by modality ( #48085 ) by @tarekziade in [ #48085 ] [VITS] Un-skip test_model_forward ( #46375 ) by @blipbyte in [ #46375 ] Apply context parallelism to the evaluation path ( #48167 ) by @qgallouedec in [ #48167 ] Fix gpt_oss runs on GPU ( #48118 ) by @tarekziade in [ #48118 ] [EsmFold2] Fix stale expected distogram logit values ( #48182 ) by @ydshieh in [ #48182 ] Fix Apr 05 integration test regressions (cuda sm_86) ( #48170 ) by @ydshieh in [ #48170 ] Add MLU support to is_flash_linear_attention_available ( #46995 ) by @atri2549 in [ #46995 ] Fix integration test expected values for cuda sm_86 (Mar 15 regressions) ( #48168 ) by @ydshieh in [ #48168 ] Fix DynamicCache reconstruction during ExecuTorch export ( #47900 ) by @eladsegal in [ #47900 ] [LLaVA] Fix pixtral integration tests for cuda sm_86 ( #48166 ) by @ydshieh in [ #48166 ] Port ESMC and ESMFold2 to Transformers ( #46419 ) by @Rocketknight1 in [ #46419 ] [Qwen2.5-Omni] Update stale expected values for cuda sm_86 ( #48164 ) by @ydshieh in [ #48164 ] [Mistral3] Fix batched integration tests: padding_side=left + update expected values ( #48161 ) by @ydshieh in [ #48161 ] Fix DeepSeek V2 default vocab size ( #48159 ) by @hmellor in [ #48159 ] [InternVL] Fix stale expected values for Llama integration tests (cuda sm_80) ( #48153 ) by @ydshieh in [ #48153 ] Retry transient network errors (RemoteDisconnected) in github_utils ( #48124 ) by @ydshieh in [ #48124 ] fix bugs for clvp model ( #47127 ) by @kaixuanliu in [ #47127 ] Always tie embeddings for LongT5 and Pop2Piano ( #47620 ) by @jiqing-feng in [ #47620 ] Use generator with seed for LengthGroupedSampler in Trainer._get_eval_sampler for deterministic eval order with per_device_eval_batch_size > 1 ( #48025 ) by @philipshurpik in [ #48025 ] Enable mlinter findings artifact for inline PR reviews ( #48117 ) by @ydshieh in [ #48117 ] Fix CpmAnt loading: size lm_head to vocab_size ( #48012 ) by @jiqing-feng in [ #48012 ] Delete old mlinter review comments before posting new ones ( #48107 ) by @ydshieh in [ #48107 ] [Video] Warn and return all frames when num_frames exceeds total_num_frames ( #48074 ) by @carlszk in [ #48074 ] Accept artifact dir as argument in post_mlinter_review.py ( #48106 ) by @ydshieh in [ #48106 ] Fix mlinter artifact path ( #48088 ) by @ydshieh in [ #48088 ] [Video] Fix convert_to_rgb channel slicing and alpha blending for RGBA videos ( #48053 ) by @ in [ #48053 ] [PE] Skip test_sdpa_can_dispatch_on_flash for TimmWrapper-backed models ( #48064 ) by @ydshieh in [ #48064 ] fix(pipeline): preserve model.generation_config precedence over pipeline defaults ( #47752 ) ( #47953 ) by @nithin42 in [ #47953 ] [Gemma3] Update integration test expected values for A10G ( #48036 ) by @ydshieh in [ #48036 ] Fallback from 'lanczos' to 'bicubic' when on cuda ( #48026 ) by @zucchini-nlp in [ #48026 ] [serge] Fix 2 integration tests regressed by commit b9090ae (PR #47096 ) ( #48060 ) by @sergereview[bot] in [ #48060 ] [Fix] Small FA-related test failures in CB ( #47341 ) by @remi-or in [ #47341 ] fix: honor empty processor_kwargs={} in multimodal pipelines ( #48044 ) by @ in [ #48044 ] [CircleCI] Enable CI for private forks, no-op for public repo ( #48056 ) by @ydshieh in [ #48056 ] Support BatchFeature in length-grouped samplers ( #48034 ) by @qgallouedec in [ #48034 ] Fix EOS for candidate generators ( #47931 ) by @ in [ #47931 ] [Gemma3n] Update integration test expected values for A10G + torch 2.13 ( #48035 ) by @ydshieh in [ #48035 ] fix failed test cases for muse_glimmer ( #48011 ) by @kaixuanliu in [ #48011 ] Cohere compass tests ( #47895 ) by @zucchini-nlp in [ #47895 ] docs: use relative paths for README language menus and add fa/ro entries ( #47777 ) by @Priyans-Lathiya in [ #47777 ] Moving mlinter to 0.1.4 ( #47918 ) by @tarekziade in [ #47918 ] [MoE] Fix Blackwell GPU crash with torch._grouped_mm on torch <= 2.8 ( #48014 ) by @ in [ #48014 ] [docs] Muse Glimmer ( #47882 ) by @stevhliu in [ #47882 ] [Florence2] Fix two integration test failures caused by torch 2.13 and auto-dtype ( #48031 ) by @ydshieh in [ #48031 ] Proper separation of tests ( #47943 ) by @zucchini-nlp in [ #47943 ] unpin pytest in the examples_torch deps ( #48023 ) by @tarekziade in [ #48023 ] [CI] Fix startup failure in pr_build_doc_with_comment workflow by adding missing get-pr-number dependency ( #47971 ) by @ in [ #47971 ] Let the GPU verify caller turn on the memory probe ( #48001 ) by @tarekziade in [ #48001 ] Fix MTP config when mlp_layer_types is absent ( #48015 ) by @Cyrilvallez in [ #48015 ] Remove duplicate block_sparse_moe assignment in GraniteMoeDecoderLayer ( #47876 ) by @Aman2394 in [ #47876 ] [ModernVBERT] Fix integration test checkpoint (404 since April) ( #48009 ) by @ydshieh in [ #48009 ] Fix cropping ( #48006 ) by @Cyrilvallez in [ #48006 ] Fix DFlash candidate token device mismatch with device_map=\"auto\" ( #47877 ) by @sywangyi in [ #47877 ] 🔴 Allow tokenizers 0.23.1 ( #46381 ) by @ArthurZucker in [ #46381 ] [Whisper] Fix batch decode_with_timestamps in WhisperTokenizer.decode() ( #47997 ) by @ydshieh in [ #47997 ] [OLMoE] Update expected logits for A10G and add torch.no_grad() ( #47989 ) by @ydshieh in [ #47989 ] [OLMo] Fix OOM in logits tests by adding torch.no_grad() ( #47986 ) by @ydshieh in [ #47986 ] [AXK1] Fix expected logits for CUDA A10G ( #47980 ) by @ydshieh in [ #47980 ] [Gemma] Update expected values for A10G ( #47976 ) by @ydshieh in [ #47976 ] Fix gemma4 video to device ( #47896 ) by @guarin in [ #47896 ] Remove stale (None, None) fallback in qwen2_5_vl batch_different_resolutions test ( #47972 ) by @ydshieh in [ #47972 ] Potential fix for code scanning alert no. 267: Artifact poisoning ( #47949 ) by @tarekziade in [ #47949 ] Fix GatedDeltaNet A_log dtype to prevent -inf under bfloat16 init ( #47944 ) by @Nkluge-correa in [ #47944 ] [serge] Fix 2 integration tests for model got_ocr2 failing with other (other (2)) ( #47937 ) by @sergereview[bot] in [ #47937 ] Fix Jinja block endings in CHAT WITH MODELS' Writing a chat template … ( #47960 ) by @ak1for2business-prog in [ #47960 ] [tests] Fix expected output for Qwen2.5-VL batch_wo_image on CUDA ( #47968 ) by @ydshieh in [ #47968 ] [serge] Fix 2 integration tests for model opt failing with other (other (2)) ( #47909 ) by @sergereview[bot] in [ #47909 ] Scan a diff in trufflehog, not the whole repo history ( #47945 ) by @tarekziade in [ #47945 ] [serge] Fix 2 integration tests for model vivit failing with output_mismatch (tensor values differ (2)) ( #47566 ) by @sergereview[bot] in [ #47566 ] Fix Gemma sliding_window being halved on every config save/reload ( #47940 ) by @Bluear7878 in [ #47940 ] docs: fix incorrect PEFT anchor link in fine-tuning section ( #47927 ) by @dsulot in [ #47927 ] CI: add vllm-test-init and vllm-test-transformers jobs on dedicated runners ( #47934 ) by @ydshieh in [ #47934 ] Make muse glimmer exportable ( #47871 ) by @IlyasMoutawwakil in [ #47871 ] docs: add installation instructions for NVIDIA Spark (ARM64) devices ( #47906 ) by @mfuntowicz in [ #47906 ] [CohereCompass] Minor docs fixes ( #47903 ) by @calpt in [ #47903 ] fix: correct checkpoints, config annotations, and create_dummy_models improvements ( #47902 ) by @ydshieh in [ #47902 ] Transform paths and repeat joining for response parsing ( #47648 ) by @Rocketknight1 in [ #47648 ] [docs] Update toctree ( #47781 ) by @stevhliu in [ #47781 ] Add CI_CPU_MEMORY_LIMIT_GB to check_failed_tests workflow ( #47884 ) by @ydshieh in [ #47884 ] Use tiny Hub checkpoint in Qwen3ASR processor test ( #47833 ) by @ydshieh in [ #47833 ] Update AutoRound XPU/CPU backend ( #47826 ) by @yiliu30 in [ #47826 ] docs(tests): fix typos in test comments ( #47859 ) by @zhaoxinyi02 in [ #47859 ] Update version post release ( #47870 ) by @Cyrilvallez in [ #47870 ] Significant community contributions The following contributors have made significant changes to the library over the last release: @tarekziade CI: gate the hunyuan-moe slow test ( #48330 ) CI: fix muse OOMs ( #48284 ) ignore mlinter ci file ( #48267 ) Post two CI badges on a PR: CPU PR CI and GPU run-slow ( #48190 ) Assign a reviewer even when a codeowner has left, and route models by modality ( #48085 ) Fix gpt_oss runs on GPU ( #48118 ) Moving mlinter to 0.1.4 ( #47918 ) unpin pytest in the examples_torch deps ( #48023 ) Let the GPU verify caller turn on the memory probe ( #48001 ) Potential fix for code scanning alert no. 267: Artifact poisoning ( #47949 ) Scan a diff in trufflehog, not the whole repo history ( #47945 ) @dkrisman Add an opt-in per-frame pixel cap (cap_pixels_per_frame) to the Qwen3-VL video processor ( #48071 ) @jiqing-feng Cpmant fix use cache ( #48013 ) [xcodec2] Fix flex attention and flash dispatch tests ( #48244 ) Always tie embeddings for LongT5 and Pop2Piano ( #47620 ) Fix CpmAnt loading: size lm_head to vocab_size ( #48012 ) Fix Qwen2.5-Omni / Qwen3-Omni-MoE generation with a compilable cache ( #47872 ) Bump default flash-attn2 hub kernel version to v3 ( #47863 ) Declare sdpa support in TimmWrapper ( #47939 ) @ydshieh Fix AutoTokenizer returning TokenizersBackend for DeepSeek-R1-Distill-Qwen models ( #48211 ) [Gemma4] Investigate flaky test_generation_beyond_sliding_window_1_eager ( #48236 ) [Gemma4] Fix stale expected values in integration tests ( #48233 ) [TableTransformer, PI0] Fix stale expected values and OOM in integration tests ( #48198 ) Fix stale expected values in integration tests (cuda sm_86 / Aug04 regressions) ( #48171 ) [EsmFold2] Fix stale expected distogram logit values ( #48182 ) Fix Apr 05 integration test regressions (cuda sm_86) ( #48170 ) Fix integration test expected values for cuda sm_86 (Mar 15 regressions) ( #48168 ) [LLaVA] Fix pixtral integration tests for cuda sm_86 ( #48166 ) [Qwen2.5-Omni] Update stale expected values for cuda sm_86 ( #48164 ) [Mistral3] Fix batched integration tests: padding_side=left + update expected values ( #48161 ) [InternVL] Fix stale expected values for Llama integration tests (cuda sm_80) ( #48153 ) Retry transient network errors (RemoteDisconnected) in github_utils ( #48124 ) [Whisper] Fix speculative decoding: preserve cleared suppress tokens through super().generate() ( #48108 ) Enable mlinter findings artifact for inline PR reviews ( #48117 ) Delete old mlinter review comments before posting new ones ( #48107 ) Accept artifact dir as argument in post_mlinter_review.py ( #48106 ) Fix mlinter artifact path ( #48088 ) [Whisper] Fix decoder position IDs for left-padded batches in longform generation ( #48028 ) [PE] Skip test_sdpa_can_dispatch_on_flash for TimmWrapper-backed models ( #48064 ) [Gemma3] Update integration test expected values for A10G ( #48036 ) [CircleCI] Enable CI for private forks, no-op for public repo ( #48056 ) [Gemma3n] Update integration test expected values for A10G + torch 2.13 ( #48035 ) [Florence2] Fix two integration test failures caused by torch 2.13 and auto-dtype ( #48031 ) [ModernVBERT] Fix integration test checkpoint (404 since April) ( #48009 ) [Whisper] Fix speculative decoding: UnboundLocalError, cache corruption, and speed regression ( #48000 ) [Whisper] Fix batch decode_with_timestamps in WhisperTokenizer.decode() ( #47997 ) [Whisper] Fix integration test failures on A10G (dtype, stale values, API changes) ( #47995 ) [DeepSeekV2] Fix integration tests OOM: use device_map=auto instead of 8-bit quantization ( #47991 ) [OLMoE] Update expected logits for A10G and add torch.no_grad() ( #47989 ) [OLMo] Fix OOM in logits tests by adding torch.no_grad() ( #47986 ) [GPTNeoX] Fix post_processor not overridden when loading from pretrained (OLMo garbage generation) ( #47988 ) [AXK1] Fix expected logits for CUDA A10G ( #47980 ) [Gemma] Update expected values for A10G ( #47976 ) [emu3] 🦮 Black Labrador is back! Fix image generation broken since #37033 ( #47948 ) Remove stale (None, None) fallback in qwen2_5_vl batch_different_resolutions test ( #47972 ) [tests] Fix expected output for Qwen2.5-VL batch_wo_image on CUDA ( #47968 ) CI: add vllm-test-init and vllm-test-transformers jobs on dedicated runners ( #47934 ) fix: correct checkpoints, config annotations, and create_dummy_models improvements ( #47902 ) Add CI_CPU_MEMORY_LIMIT_GB to check_failed_tests workflow ( #47884 ) Use tiny Hub checkpoint in Qwen3ASR processor test ( #47833 ) @eustlb gs ( #48288 ) @eladsegal Fix DynamicCache reconstruction during ExecuTorch export ( #47900 ) Support per-layer cache configuration and attention-mask selection ( #47901 ) @YangKai0616 🚨[wav2vec2] Support attn_implementation=sdpa dispatch ( #46196 ) @drbh feat: add nvfp4 quantization ( #47883 ) @Priyans-Lathiya docs: use relative paths for README language menus and add fa/ro entries ( #47777 ) @itazap [new model] step 3.7 ( #46658 ) @calpt [CohereCompass] Minor docs fixes ( #47903 ) Add CohereCompass modeling ( #47878 )",
  "tags": [
    "Hugging Face",
    "huggingface.releases",
    "ai",
    "ml",
    "transformers",
    "open-source"
  ]
}

Get Your Free API Key

Sign up to access the full changelog API. All public sources are free — no credit card required.

Sign Up Free →

Tags:

aimltransformersopen-source

Related Sources

Favicon

 

  
  
Favicon

 

  
  
Favicon

 

  
  

Share: