transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-03 12:50:06 +06:00

Author	SHA1	Message	Date
Younes Belkada	6829936ee0	[MODEL] Add Falcon H1 (#38249 ) * Create push-important-models.yml * feat: add falcon-h1 * fixup * address comment * fix * fix copies * fix copies * fix * fix * fix * fix * fix copies * fix * fix copies * fix test import to at least trigget the cis * yups * update * fix make fix copies * fix inits? * fix style * skip annoying test * add integration test for Falcon H1 * fix copies * fix --------- Co-authored-by: Arthur Zucker <arthur.zucker@gmail.com> Co-authored-by: dhia.rhaiem <dhia.rhaiem@tii.ae>	2025-05-21 10:43:11 +02:00
Arthur	e288ee00d8	tp plan should not be NONE (#38255 ) * accept custom device_mesh * fix device_map * assert that num_heads % tp_size == 0 * todo. * ReplicateParallel * handle tied weights * handle dtensor in save_pretrained with safe_serialization * tp test works * doesnt work * fix shard_and_distribute_module's rank should be local_rank * tp=4 is correct * dp+tp is broken * todo allreduce with dtensors on another dim is annoying * workaround to sync dp grads when using dtensors * loading a checkpoint works * wandb and compare losses with different tp/dp * cleaning * cleaning * . * . * logs * CP2 DP2 no mask works after commenting attn_mask and is_causal from scaled_dot_product_attention * DP=2 TP=2 now works even with tied embeddings * model.parameters() and model.module.parameters() are empty.. * reformat sanity_check_tensor_sync * set atol=1e-4 for CP to pass * try populate _parameters from named_modules * refactors TP2 DP2 works CP2 DP2 works * is_causal=True and pack sequences, no attn mask, and preshuffle dataset * fix packing * CP=4 doesn't work * fix labels and position_ids for CP * DP CP works with transformers 🥳🥳🥳 * refactor * add example cp * fixup * revert sdpa changes * example cleared * add CP, DP to the mesh init * nit * clean * use `ALL_PARALLEL_STYLES` * style * FSDP works * log on 1 rank * . * fix? * FSDP1 also has .parameters() bug * reported gradnorm when using FSDP1 is wrong, but loss is correct so it's okay * . * style and fixup * move stuff around * fix tests * style * let's make it a check * add missing licences * warning should be an info * tp plan should not be NONE * test all * god damn it * test all --------- Co-authored-by: nouamanetazi <nouamane98@gmail.com>	2025-05-21 10:22:38 +02:00
Lysandre Debut	711d78d104	Revert parallelism temporarily (#38240 ) Some checks are pending Self-hosted runner (benchmark) / Benchmark (aws-g5-4xlarge-cache) (push) Waiting to run Details Build documentation / build (push) Waiting to run Details Slow tests on important models (on Push - A10) / Get all modified files (push) Waiting to run Details Slow tests on important models (on Push - A10) / Slow & FA2 tests (push) Blocked by required conditions Details Self-hosted runner (push-caller) / Check if setup was changed (push) Waiting to run Details Self-hosted runner (push-caller) / build-docker-containers (push) Blocked by required conditions Details Self-hosted runner (push-caller) / Trigger Push CI (push) Blocked by required conditions Details Secret Leaks / trufflehog (push) Waiting to run Details Update Transformers metadata / build_and_package (push) Waiting to run Details * Revert "Protect ParallelInterface" This reverts commit `cb513e35f9`. * Revert "parallelism goes brrr (#37877)" This reverts commit `1c2f36b480`. * Empty commit	2025-05-20 22:43:04 +02:00
Yih-Dar	feec294dea	CI reporting improvements (#38230 ) update Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2025-05-20 19:34:58 +02:00
Lysandre	cb513e35f9	Protect ParallelInterface	2025-05-20 18:27:50 +02:00
Lysandre	f4ef41c45e	v4.53.0.dev0	2025-05-20 18:12:56 +02:00
Raushan Turganbay	f834d368f6	[gemma3] fix bidirectional attention mask (#38080 ) * fix attn mask * attn viz doesn't show yello cubes between images * bucketize made it hard with different number of crops * fixup	2025-05-20 17:35:04 +02:00
Raushan Turganbay	2edb0e4b4d	[mllama] fix loading and inference (#38223 ) fix loading	2025-05-20 17:34:55 +02:00
Garrett Goon	390f153469	Add padding-free to bamba (#35861 ) * add seq_idx and fa kwargs * update tests * docs and grad ckpt support * fmt * better names * test_raise_missing_padding_free_kwarg_errs * + seq_idx in doc strings * padding free training docs * add link to pr plots * raise err on attn_mask with padding free * rm raising missing padding free err test * BambaFlashAttentionKwargs * run modular util for modular_granitemoehybrid.py	2025-05-20 17:13:59 +02:00
Mohamed Mekkouri	2a79471318	Fixing Bitnet after use_rms_norm introduction (#38229 ) * fix * make style	2025-05-20 17:13:21 +02:00
Bojun Feng	9661896083	Enable Quantize KV Cache for Mistral Model (#35042 ) fix #35041	2025-05-20 16:50:26 +02:00
Nouamane Tazi	1c2f36b480	parallelism goes brrr (#37877 ) * accept custom device_mesh * fix device_map * assert that num_heads % tp_size == 0 * todo. * ReplicateParallel * handle tied weights * handle dtensor in save_pretrained with safe_serialization * tp test works * doesnt work * fix shard_and_distribute_module's rank should be local_rank * tp=4 is correct * dp+tp is broken * todo allreduce with dtensors on another dim is annoying * workaround to sync dp grads when using dtensors * loading a checkpoint works * wandb and compare losses with different tp/dp * cleaning * cleaning * . * . * logs * CP2 DP2 no mask works after commenting attn_mask and is_causal from scaled_dot_product_attention * DP=2 TP=2 now works even with tied embeddings * model.parameters() and model.module.parameters() are empty.. * reformat sanity_check_tensor_sync * set atol=1e-4 for CP to pass * try populate _parameters from named_modules * refactors TP2 DP2 works CP2 DP2 works * is_causal=True and pack sequences, no attn mask, and preshuffle dataset * fix packing * CP=4 doesn't work * fix labels and position_ids for CP * DP CP works with transformers 🥳🥳🥳 * refactor * add example cp * fixup * revert sdpa changes * example cleared * add CP, DP to the mesh init * nit * clean * use `ALL_PARALLEL_STYLES` * style * FSDP works * log on 1 rank * . * fix? * FSDP1 also has .parameters() bug * reported gradnorm when using FSDP1 is wrong, but loss is correct so it's okay * . * style and fixup * move stuff around * fix tests * style * let's make it a check * warning should be an info --------- Co-authored-by: Arthur Zucker <arthur.zucker@gmail.com>	2025-05-20 16:22:52 +02:00
Cyril Vallez	b591d925be	Fix Llama4 (#38222 ) Update modeling_llama4.py	2025-05-20 16:00:46 +02:00
ivarflakstad	3f0b7d0fac	Mamba2 remove unecessary test parameterization (#38227 )	2025-05-20 13:54:04 +00:00
Pablo Montalvo	9cde2f5d42	Minor llama4 fixes (#38123 ) * fix wrong scaling value/default Cache init * style * fix various issues on integration tests * change expected outputs * fixup * fix config access * protect default scaling	2025-05-20 13:15:54 +00:00
James Niken	856f034f45	fix dead flax links modeling_flax_pytorch_utils.py (#38212 )	2025-05-20 13:03:41 +00:00
Marc Sun	bb3c6426d8	Make `train_dataset` attribute in `_get_train_sampler` optional (#38226 ) make it optional	2025-05-20 12:59:53 +00:00
Boian Petkantchin	2ad152f84c	In Llama4 fix wrongly inverted causal attention mask when using SDPA implementation (#38094 ) When preparing the causal attention mask at this point the mask comes in as a float tensor with min value as a masked value. It is not correct to convert it to bool and treat it as a bool mask as this inverts the mask. `torch.nn.functional.scaled_dot_product_attention` expects that a masked value is `False`. I suspect that the `sdpa` implementation variant may not have been thoroughly tested and that is why this error was not caught earlier. Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2025-05-20 14:47:59 +02:00
ivarflakstad	de70c8426e	Disable torchscript tests for AriaForConditionalGenerationModelTest (#38225 ) Co-authored-by: Yih-Dar <2521628+ydshieh@users.noreply.github.com>	2025-05-20 14:37:55 +02:00
brenoca	8ea61c4530	Add support to Marimo Notebooks and Enverge.ai (#38210 ) * Add support to Marimo notebooks * Consice logic * Simplify logic * Ruff fixes	2025-05-20 12:26:34 +00:00
Manuel de Prada Corral	d34e21e7dd	New cache tests and refactored Hybrid Cache (#37972 )	2025-05-20 12:46:13 +02:00
Matthew Hoffman	183fb3637c	Add `Llama4TextModel` to `AutoModel` mapping (#38162 ) Add Llama4TextModel to AutoModel mapping using Llama4TextConfig on AutoModel.from_config raises a ValueError when it is expected to instantiate a Llama4TextModel	2025-05-20 10:01:00 +00:00
Titus	f022bf9322	Remove trust_remote_code=True tests from bnb quantization tests (MPT now integrated) (#38206 ) bnb quant tests: remove obsolete trust_remote_code test The MPT model is now natively integrated in Transformers and no longer requires trust_remote_code=True. This removes the failing test_get_keys_to_not_convert_trust_remote_code and related usage, which depended on remote code and caused CI issues due to missing dependencies (e.g., triton_pre_mlir).	2025-05-20 11:43:11 +02:00
Raushan Turganbay	0a52bd2403	[fix] sliding window attention mask (#38045 ) * fix sliding attn * make style * Update tests/test_modeling_common.py Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com> * no a second throught, should default to `True` fo BC --------- Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>	2025-05-20 09:32:19 +00:00
Yong Hoon Shin	555715f418	Fix broken example generation script for Llama3 (#38062 ) Fix broken example generation script for llama3	2025-05-20 10:53:43 +02:00
Matej Sirovatka	7a611f0afd	Fix: make docs work better with doc builder (#38213 )	2025-05-20 08:23:03 +00:00
Yao Matrix	3bd1c20149	enable misc cases on XPU & use device agnostic APIs for cases in tests (#38192 ) * use device agnostic APIs in tests Signed-off-by: Matrix Yao <matrix.yao@intel.com> * more Signed-off-by: Matrix Yao <matrix.yao@intel.com> * fix style Signed-off-by: Matrix Yao <matrix.yao@intel.com> * add reset_peak_memory_stats API Signed-off-by: YAO Matrix <matrix.yao@intel.com> * update --------- Signed-off-by: Matrix Yao <matrix.yao@intel.com> Signed-off-by: YAO Matrix <matrix.yao@intel.com> Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2025-05-20 10:09:01 +02:00
shawn	dbc4b91db4	Qwen2.5-Omni: Update modeling_qwen2_5_omni.py to fix error when loading quantized weights with AutoAWQ. (#38013 ) * Update modular_qwen2_5_omni.py fix the error when loading quantized model by AuotAWQ. * Update modeling_qwen2_5_omni.py sync code to modular_qwen2_5_omni.py	2025-05-20 09:53:51 +02:00
Matej Sirovatka	46a4b7c909	Feat: save_pretrained for tensor parallel (and other parallelisms) models (#37919 ) * tmp: initial save pretrained with dtensors * Feat: add correctness tests * Refactor: version checks * Temp: 1:1 checkpoint llama4 * refactor * Tests * Feat: works * Style * Feat: version checks + minor fixes * Style * Fix: version checks in tests * Feat: move more stuff into tensor_parallel.py	2025-05-19 18:16:21 +00:00
Fanli Lin	9ecee14378	[doc] fix bugs in `how_to_hack_models.md` (#38198 ) fix several bugs	2025-05-19 10:37:54 -07:00
Nanji Huaji	f524439cc5	Translating model_doc/bert.md to Chinese (#37806 ) * Translated model_doc/bert.md * Revise grammatical errors * Changed _toctree.yml * Revise some errors	2025-05-19 10:14:57 -07:00
Matej Sirovatka	6e738411e1	Tensor parallel docs (#38178 ) * Feat: initial docs * Feat: update doc * Final typos/changes * Refactor: reorder top to bottom.	2025-05-19 17:05:01 +00:00
Joao Gante	9c500015c5	🚨🚨🚨 [pipelines] update defaults in pipelines that can `generate` (#38129 ) * pipeline generation defaults * add max_new_tokens=20 in test pipelines * pop all kwargs that are used to parameterize generation config * add class attr that tell us whether a pipeline calls generate * tmp commit * pt text gen pipeline tests passing * remove failing tf tests * fix text gen pipeline mixin test corner case * update text_to_audio pipeline tests * trigger tests * a few more tests * skips * some more audio tests * not slow * broken * lower severity of generation mode errors * fix all asr pipeline tests * nit * skip * image to text pipeline tests * text2test pipeline * last pipelines * fix flaky * PR comments * handle generate attrs more carefully in models that cant generate * same as above	2025-05-19 18:02:06 +01:00
Joao Gante	6f9da7649f	[image-text-to-text pipeline] Accept a chat as a positional arg (#38204 ) accept chat as a positional arg	2025-05-19 17:26:09 +01:00
NielsRogge	7c9b0ca08c	[SAM-HQ] Update names in the docs (#38058 ) Update names	2025-05-19 09:21:14 -07:00
Daize Dong	04282a9ef5	Remove Deprecated `verbose` arg in LayerWiseDummyScheduler (#38197 ) Remove Deprecated args in LayerWiseDummyScheduler	2025-05-19 13:49:11 +00:00
Shane A	aef12349b6	Make HF implementation match original OLMo 2 models for lower precisions (#38131 ) * Make HF implementation match OLMo models for lower precisions * Add test of 1B logits in bfloat16 * Run make fixup	2025-05-19 15:35:23 +02:00
Fanli Lin	9644acb7cb	[docs] add Audio import (#38195 ) add Audio import	2025-05-19 13:16:35 +00:00
Fanli Lin	7d93f93f83	[docs] minor fixes in `models.md` (#38193 ) minor gix	2025-05-19 13:14:21 +00:00
Sergio Paniego Blanco	47f8578d96	Pass `eps` to `Mistral3RMSNorm` (#38026 ) Pass eps to Mistral3RMSNorm Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2025-05-19 15:09:25 +02:00
Emmanuel Ferdman	6c6302817d	Resolve Python logger warnings (#38183 ) * Resolve Python logger warnings Signed-off-by: Emmanuel Ferdman <emmanuelferdman@gmail.com> * Apply style fixes --------- Signed-off-by: Emmanuel Ferdman <emmanuelferdman@gmail.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>	2025-05-19 12:53:07 +00:00
Lysandre Debut	003deb16f1	Support for transformers explicit filename (#38152 ) * Support for transformers explicit filename * Tests * Rerun tests	2025-05-19 14:33:47 +02:00
Joao Gante	dbb9813dff	[generation] Less verbose warnings by default (#38179 ) * tmp commit (imports broken) * working version; update tests * remove line break * shorter msg * dola checks need num_beams=1; other minor PR comments * update early trainer failing on bad gen config * make fixup * test msg	2025-05-19 10:03:37 +00:00
Daize Dong	656e2eab3f	Add adam_kwargs for Apollo Optimizer (#38168 ) Add adam_kwargs for Apollo	2025-05-19 08:59:49 +00:00
Yaswanth Gali	6bb6821d93	Refactor `get_XXX_dataloader` from Trainer (#38090 ) * Remove test_dataloader * refactor	2025-05-19 10:43:27 +02:00
Joao Gante	40a493c7ed	[tests] remove `test_sdpa_equivalence` (redundant) (#37911 ) * rm test_sdpa_equivalence * make fixup --------- Co-authored-by: Yih-Dar <2521628+ydshieh@users.noreply.github.com>	2025-05-16 18:37:27 +01:00
kang sheng	ea29f61ed9	fix bug in distributed loss test (#38166 ) * fix bug in distributed loss test and change some config to pass at both 2&8 gpus * fix doc	2025-05-16 16:21:35 +00:00
Chachura Baptiste	a4389494c7	Fix import torchao.prototype.low_bit_optim since torchao v0.11 (#38174 ) * Fix ModuleNotFoundError torchao.prototype.low_bit_optim since torchao v 0.11.0 * Fix space on blank line * update torchao's AdamW4bit and AdamW8bit import for v0.11.0 * Apply style fixes --------- Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>	2025-05-16 18:02:33 +02:00
Yoni Gozlan	0ba95564b7	Add args support for fast image processors (#37018 ) * add args support to fast image processors * add comment for clarity * fix-copies * Handle child class args passed as both args or kwargs in call and preprocess functions * revert support args passed as kwargs in overwritten preprocess * fix image processor errors	2025-05-16 12:01:46 -04:00
Peter St. John	d69945e5fc	[ESM] Add flash-attention-2 backend for ESM-2 (#38023 ) * Add flash-attention-2 backend for ESM-2 Signed-off-by: Peter St. John <pstjohn@nvidia.com> * update extended_attention_mask for fa2 Signed-off-by: Peter St. John <pstjohn@nvidia.com> * add test_flash_attn_2_equivalence test Signed-off-by: Peter St. John <pstjohn@nvidia.com> --------- Signed-off-by: Peter St. John <pstjohn@nvidia.com>	2025-05-16 14:11:56 +01:00

... 3 4 5 6 7 ...

19219 Commits