transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-31 02:02:21 +06:00

Author	SHA1	Message	Date
ShunanZhu	a7feae190f	Fix remove unused parameter in docs (#35306 ) remove unused parameter in example Co-authored-by: zzzzzsa <zzzzzsaqwq@gmail.com>	2024-12-17 09:34:41 -08:00
Jacky Lee	927c3e39ec	Fix image preview in multi-GPU inference docs (#35303 ) fix: link for img	2024-12-17 09:33:50 -08:00
Jacky Lee	4302b27719	Fix typos in translated quicktour docs (#35302 ) * fix: quicktour typos * fix: one more	2024-12-17 09:32:00 -08:00
Pavel Iakubovskii	deac971c46	🚨🚨🚨 Limit backtracking in Nougat regexp (#35264 ) * Limit backtracking in regexp * Update * [run-slow] nougat * Update	2024-12-17 16:34:18 +00:00
Yih-Dar	d29a06e39a	remove `benchmark` job in `push-important-models.yml` (#35292 ) remove-bench Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2024-12-17 17:27:26 +01:00
Matt	e0ae9b5974	🚨🚨🚨 Delete conversion scripts when making release wheels (#35296 ) * Delete conversion scripts when making release wheels * make fixup * Update docstring	2024-12-17 14:18:42 +00:00
Magnus	6eb00dd2f0	Support for SDPA for SAM models (#34110 ) * feat: add support for sdpa and gradient checkpointing * fix: ruff format * fix: config sdpa * fix: sdpa layer naming convention * fix: update test_eager_matches_sdpa_inference to handle vision_hidden_states * test: skip incompatible tests and fix loading issue with sdpa - Updated tests to skip cases flash and dynamic compile. - Minor adjustment to ensure correct loading of model with sdpa for dispatch test. * style: apply Ruff formatting * ruff fix again after rebase * [run-slow] sam * [run-slow] sam * refactor: Address review comments and improve sub-config handling in SAM model tests - Added attributes for sub_configs as per PR #34410. - Enabled tests for configs, ensuring the composite model (SAM) has several sub-configs in the main config. - Added class attribute _is_composite=True to the tester class - test_sdpa_can_dispatch_composite_models added * [run-slow] sam * style: ruff * [run-slow] sam * style: ruff again ... * [run-slow] sam	2024-12-17 14:46:05 +01:00
Omar Salman	747f361da1	Add sdpa for Beit (#34941 ) * Add sdpa for Beit * Updates * [run-slow] beit * Update inference benchmarks * Update * Fix - add missed to super().forward() * Updates * Fix missing import	2024-12-17 14:44:47 +01:00
Billel Mokeddem	6c08b3b6e5	Add Falcon3 documentation (#35307 ) * Add Falcon3 documentation * Update Falcon3 documentation * Change Falcon to Falcon3 * Update docs and run make fix-copies * Add blog post and huggingface models links	2024-12-17 14:23:13 +01:00
Tony Wu	f33a0cebb3	Add ColPali to 🤗 transformers (#33736 ) * feat: run `add-new-model-like` * feat: add paligemma code with "copied from" * feat: add ColPaliProcessor * feat: add ColPaliModel * feat: add ColPaliConfig * feat: rename `ColPaliForConditionalGeneration` to `ColPaliModel` * fixup modeling colpali * fix: fix root import shortcuts * fix: fix `modeling_auto` dict * feat: comment out ColPali test file * fix: fix typos from `add-new-model-like` * feat: explicit the forward input args * feat: move everything to `modular_colpali.py` * fix: put back ColPaliProcesor * feat: add auto-generated files * fix: run `fix-copies` * fix: remove DOCStRING constants to make modular converter work * fix: fix typo + modular converter * fix: add missing imports * feat: no more errors when loading ColPaliModel * fix: remove unused args in forward + tweak doc * feat: rename `ColPaliModel` to `ColPaliForRetrieval` * fix: apply `fix-copies` * feat: add ColPaliProcessor to `modular_colpali` * fix: run make quality + make style * fix: remove duplicate line in configuration_auto * feat: make ColPaliModel inehrit from PaliGemmaForConditionalGeneration * fix: tweak and use ColPaliConfig * feat: rename `score` to `post_process_retrieval` * build: run modular formatter + make style * feat: convert colpali weights + fixes * feat: remove old weight converter file * feat: add and validate tests * feat: replace harcoded path to "vidore/colpali-v1.2-hf" in tests * fix: add bfloat16 conversion in weight converter * feat: replace pytest with unittest in modeling colpali test * feat: add sanity check for weight conversion (doesn't work yet) * feat: add shape sanity check in weigth converter * feat: make ColPaliProcessor args explicit * doc: add doc for ColPali * fix: trying to fix output mismatch * feat: tweaks * fix: ColPaliModelOutput inherits from ModelOutput instead of PaliGemmaCausalLMOutputWithPast * fix: address comments on PR * fix: adapt tests to the Hf norm * wip: try things * feat: add `__call__` method to `ColPaliProcessor` * feat: remove need for dummy image in `process_queries` * build: run new modular converter * fix: fix incorrect method override * Fix tests, processing, modular, convert * fix tokenization auto * hotfix: manually fix processor -> fixme once convert modular is fixed * fix: convert weights working * feat: rename and improve convert weight script * feat: tweaks * fest: remove `device` input for `post_process_retrieval` * refactor: remove unused `get_torch_device` * Fix all tests * docs: update ColPali model doc * wip: fix convert weights to hf * fix logging modular * docs: add acknowledgements in model doc * docs: add missing docstring to ColPaliProcessor * docs: tweak * docs: add doc for `ColPaliForRetrievalOutput.forward` * feat: add modifications from colpali-engine v0.3.2 in ColPaliProcessor * fix: fix and upload colapli hf weights * refactor: rename `post_process_retrieval` to `score_retrieval` * fix: fix wrong typing for `score_retrieval` * test: add integration test for ColPali * chore: rerun convert modular * build: fix root imports * Update docs/source/en/index.md Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com> * fix: address PR comments * wip: reduce the prediction gap in weight conversion * docs: add comment in weight conversion script * docs: add example for `ColPaliForRetrieval.forward` * tests: change dataset path to the new one in hf-internal * fix: colpali weight conversion works * test: add fine-grained check for ColPali integration test * fix: fix typos in convert weight script * docs: move input docstring in a variable * fix: remove hardcoded torch device in test * fix: run the new modular refactor * docs: fix python example for ColPali * feat: add option to choose `score_retrieval`'s output dtype and device * docs: update doc for `score_retrieval` * feat: add `patch_size` property in ColPali model * chore: run `make fix-copies` * docs: update description for ColPali cookbooks * fix: remove `ignore_index` methods * feat: remove non-transformers specific methods * feat: update `__init__.py` to new hf format * fix: fix root imports in transformers * feat: remove ColPali's inheritance from PaliGemma * Fix CI issues * nit remove prints * feat: remove ColPali config and model from `modular_colpali.py` * feat: add `ColPaliPreTrainedModel` and update modeling and configuration code * fix: fix auto-removed imports in root `__init__.py` * fix: various fixes * fix: fix `_init_weight` * temp: comment `AutoModel.from_config` for experiments * fix: add missing `output_attentions` arg in ColPali's forward * fix: fix `resize_token_embeddings` * fix: make `input_ids` optional in forward * feat: rename `projection_layer` to `embedding_proj_layer` * wip: fix convert colpali weight script * fix tests and convert weights from original repo * fix unprotected import * fix unprotected torch import * fix style * change vlm_backbone_config to vlm_config * fix unprotected import in modular this time * fix: load config from Hub + tweaks in convert weight script * docs: move example usage from model docstring to model markdown * docs: fix input docstring for ColPali's forward method * fix: use `sub_configs` for ColPaliConfig * fix: remove non-needed sanity checks in weight conversion script + tweaks * fix: fix issue with `replace_return_docstrings` in ColPali's `forward` * docs: update docstring for `ColPaliConfig` * test: change model path in ColPali test * fix: fix ColPaliConfig * fix: fix weight conversion script * test: fix expected weights for ColPali model * docs: update ColPali markdown * docs: fix minor typo in ColPaliProcessor * Fix tests and add _no_split_modules * add text_config to colpali config * [run slow] colpali * move inputs to torch_device in integration test * skip test_model_parallelism * docs: clarify quickstart snippet in ColPali's model card * docs: update ColPali's model card --------- Co-authored-by: yonigozlan <yoni.gozlan@huggingface.co> Co-authored-by: Yoni Gozlan <74535834+yonigozlan@users.noreply.github.com>	2024-12-17 11:26:43 +01:00
Arthur	a7f5479b45	fix modular order (#35297 ) * fix modular ordre * fix * style	2024-12-17 08:05:35 +01:00
UV	f5620a7634	Improved documentation of Automatic speech recognition (#35268 ) Improved documentation quality of Automatic speech recognition	2024-12-16 09:50:11 -08:00
湛露先生	eb92bc44b7	Fix wrongs in quicktour[zh] (#35272 ) Signed-off-by: zhanluxianshen <zhanluxianshen@163.com>	2024-12-16 09:23:34 -08:00
HMJ0628	886f690e76	Translating "translate perf_infer_gpu_multi.md" to Chinese (#35271 ) add "translate perf_infer_gpu_multi"	2024-12-16 09:22:35 -08:00
Jacky Lee	22834eeba1	Fix typos in Translated Audio Classification Docs (#35287 ) * fix: qwen2 model ids * fix: line * fix: more format * update: reformat * fix: doc typos	2024-12-16 08:51:32 -08:00
eustlb	9feae5fb01	[Whisper] patch float type on mps (#35295 ) * fix float type on mps * make	2024-12-16 16:52:47 +01:00
湛露先生	d5b81e1ca1	Delete redundancy for loop checks. (#35288 ) Signed-off-by: zhanluxianshen <zhanluxianshen@163.com>	2024-12-16 13:36:27 +00:00
ivarflakstad	d0f32212ed	Temporarily disable amd push ci (#35293 ) Temporarily disable amd push ci (reduce noise)	2024-12-16 14:18:50 +01:00
Mohamed Mekkouri	85eb339231	Fix : model used to test ggml conversion of Falcon-7b is incorrect (#35083 ) fixing test model	2024-12-16 13:21:44 +01:00
Raushan Turganbay	14910281a7	Blip: fix offloading and MP tests (#35239 ) * fix device map * fix offloading + model parallel test	2024-12-16 12:44:33 +01:00
Yih-Dar	66531a1ec3	Aggeregate test summary files in CircleCI workflow runs (#34989 ) * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * fix * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * try 1 * fix * fix * fix * update * fix * fix --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2024-12-16 11:06:17 +01:00
Yoni Gozlan	5615a39369	Fall back to slow image processor in ImageProcessingAuto when no fast processor available (#34785 ) * refactor image_processing_auto logic * fix fast image processor tests * Fix tests fast vit image processor * Add safeguard when use_fast True and torchvision not available * change default use_fast back to None, add warnings * remove debugging print * call get_image_processor_class_from_name once	2024-12-15 14:00:36 -05:00
French_Ball	ca03842cdc	[i18n-Chinese] Translating perf_train_cpu.md to Chinese (#35242 ) add "1"	2024-12-13 14:46:49 -08:00
Wing Lian	add53e25ff	don't use no_sync when deepspeed doesn't support it for certain zero stages (#35157 ) * don't use no_sync when deepspeed doesn't support it for certain zero stages * chore: lint * fix no_sync context for deepspeed across all zero types * chore: lint	2024-12-13 19:23:00 +01:00
Zach Mueller	7237b3ecfc	Fix FSDP no longer working (#35212 ) Fix FSDP failing	2024-12-13 19:20:51 +01:00
HMJ0628	6009642459	Translating agents_advanced.md to Chinese (#35231 ) add "translate agents_advanced"	2024-12-13 10:12:00 -08:00
UV	e94083bf90	Fixed typos in Audio Classification Documentation (#35263 ) * Fixed typos in Audio Classification Documentation * removed space in '8000 kHZ' * Changes made as per review	2024-12-13 09:43:44 -08:00
ivarflakstad	bc6ae0d55e	Update AMD docker image (rocm 6.1) (#35259 ) * Use rocm 6.3 as base amd image and add nvidia-ml-py to exclude list * Align rocm base image with torch wheels @6.1. Seems like the most stable combo	2024-12-13 15:41:03 +01:00
Yih-Dar	8096161b76	Use `rsfE` with `pytest` (#35119 ) * fix * fix --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2024-12-13 14:36:22 +01:00
Fanli Lin	bdd4201fdb	[tests] fix "Tester object has no attribute '_testMethodName'" (#34910 ) * add more cases * fix method not found in unittest Signed-off-by: Lin, Fanli <fanli.lin@intel.com> * fix more cases * add more models * add all * no unittest.case * remove for oneformer * fix style --------- Signed-off-by: Lin, Fanli <fanli.lin@intel.com>	2024-12-13 14:33:45 +01:00
nhamanasu	3d213b57fe	skip Fuyu from test_generate (#35246 ) * skip Fuyu from test_generate * make fixup, quality, repo-consistency	2024-12-13 10:12:49 +01:00
alexrs-cohere	64478c7631	Add Cohere2 model (#35224 )	2024-12-13 09:35:50 +01:00
George	e4e404fdd0	Run model as compressed/uncompressed mode (#34719 ) * draft, run model as compreszed/uncompressed mode * draft * run run_compressed=False * run_compressed as attr * set run_compressed=False using quantization_config * remove redundant line * make is_qat_trainable dependent on run_compressed status * add tests * lint * full in docstring * add decompress * comments * decompress if model is compresssed and not run_compressed * apply_quant_config logic fix -- populate statedict properly * comments * remove non compressed model * make is_compressed as property * cosmetic * run apply_quant_config for non-compressed models -- popualte scales and zeropoints * add pahtway for decompressing sparse models * typo on is_quantization_compressed * lint * fix typo	2024-12-13 08:23:31 +01:00
EricWinsorDSIT	31f9a289a6	Fix typo in chat template example (#35250 ) Fix template example typo	2024-12-12 16:53:21 -08:00
Lysandre Debut	11ba1d472c	[Init refactor] Modular changes (#35240 ) * Modular changes * Gemma * Gemma	2024-12-12 19:23:28 +01:00
Yih-Dar	a691ccb0c2	Change back to `Thread` for SF conversion (#35236 ) * fix * fix * fix --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2024-12-12 16:05:04 +01:00
Nadav Timor	e3ee49fcfb	Refactoring `AssistedCandidateGenerator` for Improved Modularity and Reusability (#35009 ) * move `TestAssistedCandidateGeneratorDifferentTokenizers` into a new testing file * refactor * NOTHING. add space to rerun github actions tests * remove it... * NOTHING. add space to rerun github actions tests * remove it... * replace: `self.prev_tokens` -> `self.prev_assistant_ids` * NOTHING. rerun CI tests * remove it * introduce `self.prev_target_ids_len` * fix style * fix style --------- Co-authored-by: Jonathan Mamou <jonathan.mamou@intel.com>	2024-12-12 15:47:05 +01:00
Reza Rahemtola	63766abe36	Support Python 3.10+ Union style in chat template type hints parsing (#35103 ) * fix(utils): Support the newest Union type in chat template * fix(utils/chat_template): Backward compatibility for the newest Union type * Update src/transformers/utils/chat_template_utils.py Co-authored-by: Matt <Rocketknight1@users.noreply.github.com> --------- Co-authored-by: Matt <Rocketknight1@users.noreply.github.com>	2024-12-12 14:07:06 +00:00
Matt	5cf11e5ab9	Fix type hints for apply_chat_template (#35216 )	2024-12-12 13:59:24 +00:00
UV	3db8e27816	Fixed typo of 'indentifier' in audio_utils.py (#35226 )	2024-12-12 13:45:04 +00:00
Vijay	a9ccdfd8e3	docs: clarify initializer_range parameter description in Idefics3VisionConfig (#35215 )	2024-12-11 11:26:18 -08:00
Yoach Lacombe	6181c6b095	Fix seamless TTS generate (#34968 ) * fix seamless tts generate * apply same fix for v2 * [run-slow] seamless_m4t, seamless_m4t_v2 * remove TODO * [run-slow] seamless_m4t, seamless_m4t_v2 * [run-slow] seamless_m4t, seamless_m4t_v2 * ignore failing test on multigpus * [run-slow] seamless_m4t, seamless_m4t_v2 * [run-slow] seamless_m4t, seamless_m4t_v2	2024-12-11 15:38:42 +01:00
Cyril Vallez	33c12e4d80	Fix CI (#35208 ) fix aria	2024-12-11 14:24:52 +01:00
Lysandre Debut	7d303efa5f	Cleanup: continue the init refactor (#35170 ) * Round 2 * Round 3	2024-12-11 14:12:34 +01:00
Pavel Iakubovskii	5fcf6286bf	Add TimmWrapper (#34564 ) * Add files * Init * Add TimmWrapperModel * Fix up * Some fixes * Fix up * Remove old file * Sort out import orders * Fix some model loading * Compatible with pipeline and trainer * Fix up * Delete test_timm_model_1/config.json * Remove accidentally commited files * Delete src/transformers/models/modeling_timm_wrapper.py * Remove empty imports; fix transformations applied * Tidy up * Add image classifcation model to special cases * Create pretrained model; enable device_map='auto' * Enable most tests; fix init order * Sort imports * [run-slow] timm_wrapper * Pass num_classes into timm.create_model * Remove train transforms from image processor * Update timm creation with pretrained=False * Fix gamma/beta issue for timm models * Fixing gamma and beta renaming for timm models * Simplify config and model creation * Remove attn_implementation diff * Fixup * Docstrings * Fix warning msg text according to test case * Fix device_map auto * Set dtype and device for pixel_values in forward * Enable output hidden states * Enable tests for hidden_states and model parallel * Remove default scriptable arg * Refactor inner model * Update timm version * Fix _find_mismatched_keys function * Change inheritance for Classification model (fix weights loading with device_map) * Minor bugfix * Disable save pretrained for image processor * Rename hook method for loaded keys correction * Rename state dict keys on save, remove `timm_model` prefix, make checkpoint compatible with `timm` * Managing num_labels <-> num_classes attributes * Enable loading checkpoints in Trainer to resume training * Update error message for output_hidden_states * Add output hidden states test * Decouple base and classification models * Add more test cases * Add save-load-to-timm test * Fix test name * Fixup * Add do_pooling * Add test for do_pooling * Fix doc * Add tests for TimmWrapperModel * Add validation for `num_classes=0` in timm config + test for DINO checkpoint * Adjust atol for test * Fix docs * dev-ci * dev-ci * Add tests for image processor * Update docs * Update init to new format * Update docs in configuration * Fix some docs in image processor * Improve docs for modeling * fix for is_timm_checkpoint * Update code examples * Fix header * Fix typehint * Increase tolerance a bit * Fix Path * Fixing model parallel tests * Disable "parallel" tests * Add comment for metadata * Refactor AutoImageProcessor for timm wrapper loading * Remove custom test_model_outputs_equivalence * Add require_timm decorator * Fix comment * Make image processor work with older timm versions and tensor input * Save config instead of whole model in image processor tests * Add docstring for `image_processor_filename` * Sanitize kwargs for timm image processor * Fix doc style * Update check for tensor input * Update normalize * Remove _load_timm_model function --------- Co-authored-by: Amy Roberts <22614925+amyeroberts@users.noreply.github.com>	2024-12-11 12:40:30 +00:00
Benjamin Bossan	bcc50cc7ce	[PEFT] Better Trainer error when prompt learning with loading best model at the end (#35087 ) Original issue: https://github.com/huggingface/peft/issues/2256 There is a potential error when using load_best_model_at_end=True with a prompt learning PEFT method. This is because Trainer uses load_adapter under the hood but with some prompt learning methods, there is an optimization on the saved model to remove parameters that are not required for inference, which in turn requires a change to the model architecture. This is why load_adapter will fail in such cases and users should instead set load_best_model_at_end=False and use PeftModel.from_pretrained. As this is not obvious, we now intercept the error and add a helpful error message.	2024-12-11 12:44:39 +01:00
Cyril Vallez	d363e71d0e	🧹 Remove deprecated RotaryEmbedding parts in the Attention layers (#34858 ) * update * style * fix missing args * remove last trace of old rope classes * remove deprecated copied from * fix copies * trigger CIs * post rebase clean-up * reverse mistral * cleanup after dropping commits * Add comment	2024-12-11 11:16:52 +01:00
Raushan Turganbay	9094b87dd4	BLIP: enable device map (#34850 ) fix device map	2024-12-11 11:03:30 +01:00
HMJ0628	10feacd88a	[i18n-<languageCode>] Translating agents.md to Chinese (#35139 ) * add "translate agents.md" * add "agents.md" * add "translate warnings" * add "totree" * add "remove transformer_agent" * add "remove transformer _agent file" --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2024-12-10 15:16:37 -08:00
John Graham Reynolds	e8508924fd	Update data collator docstrings to accurately reference Nvidia tensor core compute capability version (#35188 ) update data collator docs to reflect correct tensor core compute capability Co-authored-by: John Graham Reynolds <john.graham.reynolds@vumc.org>	2024-12-10 15:16:01 -08:00

1 2 3 4 5 ...

17600 Commits