transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-03 12:50:06 +06:00

Author	SHA1	Message	Date
Yih-Dar	ca790303f7	Pin torch == 2.6 on PR CI docker images for now (#37695 ) pin 2.6 on CircleCi images Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2025-04-23 11:47:23 +02:00
Yao Matrix	12f65ee752	enable cpu offloading for Bark on xpu (#37599 ) * enable cpu offloading of bark modeling on XPU Signed-off-by: YAO Matrix <matrix.yao@intel.com> * remove debug print Signed-off-by: YAO Matrix <matrix.yao@intel.com> * fix style Signed-off-by: YAO Matrix <matrix.yao@intel.com> * fix review comments Signed-off-by: YAO Matrix <matrix.yao@intel.com> * enhance test Signed-off-by: YAO Matrix <matrix.yao@intel.com> * update * add deprecate message Signed-off-by: YAO Matrix <matrix.yao@intel.com> * update * update * trigger CI --------- Signed-off-by: YAO Matrix <matrix.yao@intel.com> Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2025-04-23 11:37:15 +02:00
Shahruk Hossain	4f9893cbbc	fix: remove classmethod from `Qwen2_5OmniConfig.get_text_config` (#37690 ) - Since the `get_text_config` references an instance variable within the class (`self.thinker_config`), the `get_text_config` method should not be a classmethod. - Before this fix, users were getting the following error: ''' AttributeError: type object 'Qwen2_5OmniConfig' has no attribute 'thinker_config' '''	2025-04-23 09:30:57 +02:00
Vishesh-Mistry	1d9743edc2	Updated model card for mbart and mbart50 (#37619 ) * new card for mbart and mbart50 * removed comment BADGES * Update mBart overview Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * fix typo (MBart to mBart) Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * maybe fix typo * update typo and combine notes * changed notes * changed the example sentence * fixed grammatical error and removed some lines from notes example * missed one word * removed documentation resources and added some lines of example code back in notes. --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2025-04-22 12:26:47 -07:00
Jinyong Lee	fbfa1dd4db	🌐 [i18n-KO] Translated `siglip.md` to Korean (#37145 ) * docs: ko: siglip.md * feat: nmt draft * fix: manual edits * chore: Correct document title to kebab-case format Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Apply suggestions from code review Convert unnatural language to natural Korean Co-authored-by: Yijun Lee <119404328+yijun-lee@users.noreply.github.com> --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> Co-authored-by: Yijun Lee <119404328+yijun-lee@users.noreply.github.com>	2025-04-22 12:23:19 -07:00
Yao Matrix	ece79b0688	enable blip2 and emu3 cases on XPU (#37662 ) * enable blip2 and emu3 modeling cases on XPU Signed-off-by: YAO Matrix <matrix.yao@intel.com> * fix style Signed-off-by: YAO Matrix <matrix.yao@intel.com> * remove extra new line Signed-off-by: YAO Matrix <matrix.yao@intel.com> * update --------- Signed-off-by: YAO Matrix <matrix.yao@intel.com> Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2025-04-22 18:37:09 +02:00
Ken J	ca4c114dc4	Add counters for dataset classes (#37636 ) * add counters for dataset classes * fix failed code style	2025-04-22 17:30:43 +01:00
NielsRogge	d47cdae27e	[Docs] Move models to appropriate section (#37338 ) * Move models * update --------- Co-authored-by: Yih-Dar <2521628+ydshieh@users.noreply.github.com> Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2025-04-22 18:23:14 +02:00
Deepak Sahu	dbfccd3c92	typo update in the parameter name (#37655 ) See L118 and L143 for the class attribute `hidden_dim`	2025-04-22 18:14:20 +02:00
Joao Gante	de8916dde6	[docs] only build `en` docs in push CI (#37677 )	2025-04-22 17:05:11 +01:00
Joao Gante	0f8c34b0a0	[cleanup] remove old scripts in `/scripts` 🧹 🧹 (#37676 ) * rm old files * not this one	2025-04-22 16:59:03 +01:00
Yao Matrix	6673081b21	enable 6 granite cases on xpu (#37569 ) * enable 6 granite cases on XPU Signed-off-by: YAO Matrix <matrix.yao@intel.com> * make them all pass on A100 Signed-off-by: N <matrix.yao@intel.com> * fix style Signed-off-by: YAO Matrix <matrix.yao@intel.com> * update --------- Signed-off-by: YAO Matrix <matrix.yao@intel.com> Signed-off-by: N <matrix.yao@intel.com> Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2025-04-22 17:55:02 +02:00
Yao Matrix	9167461a7d	enable mllama cases on xpu (#37644 ) * enable mllama testing on xpu Signed-off-by: YAO Matrix <matrix.yao@intel.com> * more mllama cases enabling Signed-off-by: YAO Matrix <matrix.yao@intel.com> * make cases pass on A100 Signed-off-by: N <matrix.yao@intel.com> --------- Signed-off-by: YAO Matrix <matrix.yao@intel.com> Signed-off-by: N <matrix.yao@intel.com>	2025-04-22 17:39:10 +02:00
Mohamed Mekkouri	de182ba269	Refactor bitsandbytes doc (#37668 ) * doc * torch ops * fix * nits * Update docs/source/en/quantization/bitsandbytes.md Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com> --------- Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>	2025-04-22 16:13:25 +02:00
Antonin Stefanutti	dde9b03e3b	Fix no_split_modules for Llama4 pretrained models (#37673 )	2025-04-22 16:05:12 +02:00
Marc Sun	9481e9e9f1	Fix autoround docs (#37675 ) * fix * empty	2025-04-22 15:33:13 +02:00
Mohamed Mekkouri	38c406844e	Fixing quantization tests (#37650 ) * fix * style * add capability check	2025-04-22 13:59:57 +02:00
Wenhua Cheng	b3492ff9f7	Add AutoRound quantization support (#37393 ) * add auto-round support * Update src/transformers/quantizers/auto.py Co-authored-by: Ilyas Moutawwakil <57442720+IlyasMoutawwakil@users.noreply.github.com> * fix style issue Signed-off-by: wenhuach <wenhuach87@gmail.com> * tiny change * tiny change * refine ut and doc * revert unnecessary change * tiny change * try to fix style issue * try to fix style issue * try to fix style issue * try to fix style issue * try to fix style issue * try to fix style issue * try to fix style issue * fix doc issue * Update tests/quantization/autoround/test_auto_round.py * fix comments * Update tests/quantization/autoround/test_auto_round.py Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com> * Update tests/quantization/autoround/test_auto_round.py Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com> * update doc * Update src/transformers/quantizers/quantizer_auto_round.py Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com> * update * update * fix * try to fix style issue * Update src/transformers/quantizers/auto.py Co-authored-by: Mohamed Mekkouri <93391238+MekkCyber@users.noreply.github.com> * Update docs/source/en/quantization/auto_round.md Co-authored-by: Mohamed Mekkouri <93391238+MekkCyber@users.noreply.github.com> * Update docs/source/en/quantization/auto_round.md Co-authored-by: Mohamed Mekkouri <93391238+MekkCyber@users.noreply.github.com> * Update docs/source/en/quantization/auto_round.md Co-authored-by: Mohamed Mekkouri <93391238+MekkCyber@users.noreply.github.com> * update * fix style issue * update doc * update doc * Refine the doc * refine doc * revert one change * set sym to True by default * Enhance the unit test's robustness. * update * add torch dtype * tiny change * add awq convert test * fix typo * update * fix packing format issue * use one gpu --------- Signed-off-by: wenhuach <wenhuach87@gmail.com> Co-authored-by: Ilyas Moutawwakil <57442720+IlyasMoutawwakil@users.noreply.github.com> Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com> Co-authored-by: Mohamed Mekkouri <93391238+MekkCyber@users.noreply.github.com> Co-authored-by: Shen, Haihao <haihao.shen@intel.com>	2025-04-22 13:56:54 +02:00
Cyril Vallez	9608908639	Correct warm-up with fp8 (#37670 ) * start clean warmup for quantizers * style --------- Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>	2025-04-22 13:12:49 +02:00
Cyril Vallez	6614209b96	Fix duplicated weights in fp8 quantization (#37667 ) * fix fp8 * Update quantizer_finegrained_fp8.py * fix circular import * Update quantizer_finegrained_fp8.py	2025-04-22 13:12:27 +02:00
Raushan Turganbay	dcf6df5b0d	[qwen-omni] fix training (#37517 ) * fix * add text config * fixup * fix docs	2025-04-22 12:36:07 +02:00
Pavel Iakubovskii	9167fadab9	Introduce GradientCheckpointingLayer (#37223 ) * GradientCheckpointingLayer * trigger * Move GC layer to a separate file * Update import * Expose and document GC layer * Fix dummy * Apply to llama-based models * Update modulars * Update a few more models for consistency * Update glm4 * Update Janus	2025-04-22 11:33:31 +01:00
Manuel de Prada Corral	413f9bbf80	Fixes #37219 : RecurrentGemma crashes for inputs longer than sliding window length (#37613 ) * fix: RecurrentGemma crashes during inference for inputs longer than sliding window width * fix recurrentgemma tests; add long test bigger than context window	2025-04-22 12:21:16 +02:00
jeffhataws	964a1b6b7d	Fix ValueError when eval_do_concat_batches=False with examples (#37621 ) https://github.com/huggingface/transformers/issues/37593 Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>	2025-04-22 12:13:25 +02:00
Joao Gante	85665a4263	[tests] Stricter generate + compilation test -- no recompilations allowed (#37629 ) * tmp commit * stricter compilation test * trigger tests * rm todo	2025-04-22 11:12:18 +01:00
Joao Gante	362fa37da2	[test] update `test_past_key_values_format` (#37614 ) allow custom shapes	2025-04-22 11:07:34 +01:00
Manuel de Prada Corral	1cd110c6cb	Add test to ensure unknown exceptions reraising in utils/hub.py::cached_files() (#37651 ) * add test to ensure unknown exceptions are reraised in utils/hub.py::cached_files()	2025-04-22 11:38:10 +02:00
Isotr0py	c69e23455d	Support loading Gemma3 QAT GGUF models (#37649 ) * fix gemma3 qat gguf support Signed-off-by: isotr0py <2037008807@qq.com> * update test Signed-off-by: isotr0py <2037008807@qq.com> * make ruff happy Signed-off-by: isotr0py <2037008807@qq.com> --------- Signed-off-by: isotr0py <2037008807@qq.com> Co-authored-by: Mohamed Mekkouri <93391238+MekkCyber@users.noreply.github.com>	2025-04-22 11:23:17 +02:00
Jerry Zhang	7eb1107cc2	Restructure torchao quantization examples (#37592 ) * Restructure torchao quantization examples Summary: Mainly structured the examples by hardwares and then listed the recommended quantization methods for each hardware H100 GPU, A100 GPU and CPU Also added example for push_to_hub Test Plan: not required Reviewers: Subscribers: Tasks: Tags: * update * drop float8 cpu * address comments and simplify * small update * link update * minor update	2025-04-22 11:20:34 +02:00
chenin-wang	006530d285	[fix gemma] Set default value for output_attentions parameter in Gemma2 and Gemma… (#37633 ) * Set default value for output_attentions parameter in Gemma2 and Gemma3 models * update * fix * fix --------- Co-authored-by: chenin <wangzhichen@encosmart.com>	2025-04-22 11:18:17 +02:00
youngrok cha	31ea547b7a	[fix] make legacy bnb code work (#37331 ) * [fix] make legacy bnb code work * [fix] use get with default instead of getter * add test for bnb 8bit optim skip embed * [fix] style * add require annotation of bnb --------- Co-authored-by: jaycha <jaycha@ncsoft.com> Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>	2025-04-22 11:17:29 +02:00
Kero Liang	5f791281c3	Fix Qwen2.5-Omni get_chunked_index chunking functionality (#37631 ) * fix: qwen2.5 omni modular get_rope_index * test: add test for qwen2.5 omni rope index (video with audio input) * style * expected_position_ids readability * fix: use spatial_merge_size = 1 in unit test	2025-04-22 11:15:37 +02:00
JihadHammoud02	fee1190601	Refactor phi doc (#37583 ) * Added documentation for phi model * Update phi.md * Update phi.md * Update phi.md * Update docs/source/en/model_doc/phi.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/en/model_doc/phi.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/en/model_doc/phi.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/en/model_doc/phi.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Updated model card * Update phi.md * Update phi.md * Update phi.md * Update docs/source/en/model_doc/phi.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> --------- Co-authored-by: Jihad <jihadhammoud_@hotmail.com> Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2025-04-21 10:31:04 -07:00
JihadHammoud02	b2db54f66b	Update longformer.md (#37622 ) * Update longformer.md * Update longformer.md * Update docs/source/en/model_doc/longformer.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/en/model_doc/longformer.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update longformer.md --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2025-04-21 10:30:51 -07:00
Manuel de Prada Corral	2c60a442f3	fix link in kv_cache.md (#37652 ) fix typo in kv_cache.md	2025-04-21 09:01:11 -07:00
Alex Brooks	a42ba80fa5	Allow Exclusion of Input IDs from RepetitionPenaltyLogitsProcessor (#37625 ) * Allow exclusion of input IDs for repetition penalty * Add logit proc tests for rep penalty exclusion * Expose rep pen flag through generate * Only slice if needed * keep current rep pen default behavior * Revert exposing reppen changes through generate * Fix test arg * Update src/transformers/generation/logits_process.py Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com> * Rename to rep penalty kwarg * Add custom repetition penalty processor example * Validate prompt_ignore_length --------- Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>	2025-04-21 15:46:05 +01:00
Lysandre Debut	1077603410	Remove torchvision requirement from AutoImageProcessor (#37457 )	2025-04-21 14:59:33 +02:00
Joao Gante	1930e750e4	[kernels] use original forward at compile time (#37604 )	2025-04-21 13:22:47 +01:00
Yoni Gozlan	6daa3eeba5	Fix InternVL attention when using qk_norm (38B and 78B) (#37620 ) * fix internvlvision attention when using qk_norm * nit * modular	2025-04-19 21:39:08 +02:00
saswatmeher	27a25bee4f	chore: update model card for SigLIP (#37585 ) * edit siglip model card * fix syntax * Update docs/source/en/model_doc/siglip.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/en/model_doc/siglip.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/en/model_doc/siglip.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/en/model_doc/siglip.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/en/model_doc/siglip.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/en/model_doc/siglip.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * address comments --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2025-04-18 13:30:41 -07:00
Xiaojian Ma	e1f379bb09	Fixing the example in generation strategy doc (#37598 ) Update generation_strategies.md The prompt text shown in the example does not match what is inside the generated output. As the generated output always include the prompt, the correct prompt should be "Hugging Face is an open-source company".	2025-04-18 12:50:17 -07:00
Pavel Iakubovskii	4f58fc9c82	Deprecate modeling_utils.py classes (#37298 ) * Move utils classes into models * Add deprecation warnings * Remove from docs * Update config attributes check	2025-04-18 18:47:34 +01:00
Yoni Gozlan	a245011252	Add InternVL (2.5 MPO) (#35968 ) * initial commit * add convert internvl * add first end-to-end working internvl * nit prompt and image proc * add working chat template * add conversion llama-based models * add tests * pass all tests * fix isort * fix modular after main merge * add video processing for internvl * add support for interlaced images and videos * Remove processing and config from modular, add more tests * add llama model tests * Modify processor for compatibility with refactored got ocr image processor * add comments in processor * Add docs and nits * change video processing to use custom sample_indices_fn * rebase and fix tests * add processor tests * Add changes Raushan review * Use the new attention interface for the vision model * nits * add support for custom video_load_backend * remove mention to InternVLTokenizer * refactor vision model to simplify logic * refactor processor for better readibility * fix copies * fix require av processor test * refactor internVL vision * Update processor and fix processing tests * fix docstring * update convert_weights for internvl3 * change image processor to fast by default * remove do_center_crop=True in convert_weights * force use_cache to True * push_to_hub before reloading * fix internVLVision for larger models * update convert weight for qk norm * fix convert_weights * fix eos_token_id in convert * update docs and integration tests * make modifs after review * fix wrong k_norm and reduce modular * change image_token_index to image_token_id * change checkpoint to OpenGVLab org * last nits * explicitely del self.num_key_value_groups * add extra special tokens	2025-04-18 18:57:33 +02:00
we1559	b0c6ff5e13	fix issue that some example with no trainer use accelerator.end_train… (#37435 ) * fix issue that some example with no trainer use accelerator.end_training in a wrong way * reformat code --------- Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>	2025-04-18 17:59:42 +02:00
Yao Matrix	6f5014ac31	fix 2 encoder_decoder issues on XPU (#37572 ) * fix 2 encoder_decoder issues on XPU Signed-off-by: YAO Matrix <matrix.yao@intel.com> * fmt --------- Signed-off-by: YAO Matrix <matrix.yao@intel.com> Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2025-04-18 17:49:24 +02:00
Raushan Turganbay	2ba6b92a6f	[VLMs] use only `xxx_token_id` for multimodal tokens (#37573 ) * use only `xxx_token_id` for multimodal tokens * update modeling files as well * fixup * why fixup doesn't fix modular docstring first? * janus, need to update configs in the hub still * last fixup	2025-04-18 17:03:39 +02:00
Pablo Montalvo	4afd3f4820	Model debugger upgrades (#37391 ) * debugging improvements * add debugging details * add more debugging details * debug more * clean up layers + output * add summary json file * cleanup * copies 👀 * remove hooks + add documentation * draft a small test, why not * respect the format (respect it) * fixup imports * nit * add tests and configurable pruning of layers	2025-04-18 16:45:54 +02:00
Joao Gante	e5ac23081e	[Gemma3] compile ✨ (#37447 )	2025-04-18 14:55:43 +01:00
Yao Matrix	a1b82563f1	enable 6 modeling cases on XPU (#37571 ) Signed-off-by: YAO Matrix <matrix.yao@intel.com>	2025-04-18 12:28:08 +02:00
Yao Matrix	3cd6627cd7	enable 6 gemma2 cases on XPU (#37564 ) Signed-off-by: YAO Matrix <matrix.yao@intel.com>	2025-04-18 12:10:34 +02:00

1 2 3 4 5 ...

18742 Commits