transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-18 20:18:24 +06:00

Author	SHA1	Message	Date
Daniel Bogdoll	de8a0b7547	Option to set 'non_blocking' for to(device) in BatchEncoding and BatchFeature (#34883 ) * Option to set 'non_blocking' for to(device) operation for performance improvements. Defaults to 'false', thus no behavioral changes. * Enabling non_blocking in to() operation of BatchFeature. * Improved docstring on utilization of non_blocking * Force non_blocking as keyword argument Co-authored-by: Pavel Iakubovskii <qubvel@gmail.com> --------- Co-authored-by: Daniel Bogdoll <dbogdoll@umich.edu> Co-authored-by: Pavel Iakubovskii <qubvel@gmail.com>	2024-12-09 11:29:04 +01:00
UV	1452dc2514	Corrected typo in agent system prompts (#35143 )	2024-12-09 10:42:23 +01:00
NielsRogge	9e420e0269	[I-JEPA] Update docs (#35148 ) Update docs	2024-12-09 10:01:31 +01:00
kang sheng	1ccca8f48c	Fix GA loss bugs and add unit test (#35121 ) * fix GA bugs and add unit test * narrow down model loss unit test diff gap * format code to make ruff happy * send num_items_in_batch argument to decoder * fix GA loss bug in BertLMHeadModel * use TinyStories-33M to narrow down diff gap * fotmat code * missing .config * avoid add extra args --------- Co-authored-by: kangsheng <kangsheng@meituan.com>	2024-12-09 09:57:41 +01:00
Pavel Iakubovskii	c8c8dffbe4	Update I-JEPA checkpoints path (#35120 ) Update checkpoints path	2024-12-06 13:42:51 +00:00
Victor Agostinelli	7f95372c62	Add feature dim attributes to BitLinear for easier PEFT integration (#34946 ) Update bitnet.py, extremely small change to allow for easier PEFT integration Co-authored-by: Mohamed Mekkouri <93391238+MekkCyber@users.noreply.github.com>	2024-12-06 13:39:45 +01:00
Aymeric Roucher	9ad4c93536	Add Aria (#34157 ) * Add Aria --------- Co-authored-by: Cyril Vallez <cyril.vallez@gmail.com> Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2024-12-06 12:17:34 +01:00
Yih-Dar	15ab310c3a	Fix private forked repo. CI (#35114 ) fix Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2024-12-06 12:03:31 +01:00
Steven Liu	98e8062df3	[docs] top_p, top_k, temperature docstrings (#35065 ) clarify	2024-12-05 11:24:51 -08:00
Jacky Lee	44f88d8ccb	[docs] Update Python version in translations (#35096 ) update: doc version	2024-12-05 11:06:54 -08:00
Lysandre	66ab300aaf	Dev version	2024-12-05 19:12:22 +01:00
Pablo Montalvo	a5bb528471	Fix signatures for processing kwargs (#35105 ) * add conversion script * remove pg2 refs * fixup style * small update * get correct scaling * add back missing bos * fix missing config keys * might revert this pos_embeddings * fixup 9b config * fix 9b * fixup 9b conversion for good + add back num_hidden_layers * add correct query scaling for 2b, 9b, 27b * fixup 27b conversion * Additional variant: 27b-896 * Use CPU for conversion to reduce GPU RAM requirements * fix causal mask generation + formatting * fix in-training causal mask generation edge case * trigger CI * update config * update config * update config * update config * update config * update config * update config * update config * update config * move conversion file to main model dir * handle multi-images + bos token * address comments for input ids * revert ci fixes * [run-slow] paligemma * fix * [run-slow] paligemma * skip end 2 end * [run-slow] paligemma --------- Co-authored-by: Pedro Cuenca <pedro@huggingface.co> Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2024-12-05 18:15:48 +01:00
Jonathan Mamou	e27465c801	Adaptive dynamic number of speculative tokens (#34156 ) * initial commit * update strategy * add tradeoff FPR TPR with cost * all probs * fix * fix * fix style * Update src/transformers/generation/configuration_utils.py shorter docstring Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com> * import guard * fix style * add is_sklearn_available condition * vectorizing to flatten the for-loop * fix style * disable adaptation for UAG * update doc * add TestAssistedCandidateGeneratorUpdateStrategy * fix style * protect import * fix style --------- Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>	2024-12-05 17:07:33 +01:00
Yih-Dar	b0a51e5cff	Fix flaky Hub CI (`test_trainer.py`) (#35062 ) * fix * Update src/transformers/testing_utils.py Co-authored-by: Lucain <lucainp@gmail.com> * fix * fix * fix * fix * fix * fix * fix * fix * check * check * check * check * check * check * Update src/transformers/testing_utils.py Co-authored-by: Lucain <lucainp@gmail.com> * Update src/transformers/testing_utils.py Co-authored-by: Lucain <lucainp@gmail.com> * check * check * check * Final space * Final adjustment --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com> Co-authored-by: Lucain <lucainp@gmail.com>	2024-12-05 17:02:27 +01:00
Arthur	a928d9c128	[`trainer`] fix the GA `model_accepts_loss_kwargs` (#34915 ) * fix * style * values * fix	2024-12-05 16:37:46 +01:00
Raushan Turganbay	e682c17e4a	BLIP: this is correct now (#35081 ) this is correct now	2024-12-05 16:30:09 +01:00
João Marcelo	50189e36a6	Add I-JEPA (#33125 ) * first draft * add IJepaEmbeddings class * fix copy-from for IJepa model * add weight conversion script * update attention class names in IJepa model * style changes * Add push_to_hub option to convert_ijepa_checkpoint function * add initial tests for I-JEPA * minor style changes to conversion script * make fixup related * rename conversion script * Add I-JEPA to sdpa docs * minor fixes * adjust conversion script * update conversion script * adjust sdpa docs * [run_slow] ijepa * [run-slow] ijepa * [run-slow] ijepa * [run-slow] ijepa * [run-slow] ijepa * [run-slow] ijepa * formatting issues * adjust modeling to modular code * add IJepaModel to objects to ignore in docstring checks * [run-slow] ijepa * fix formatting issues * add usage instruction snippet to docs * change pos encoding, add checkpoint for doc * add verify logits for all models * [run-slow] ijepa * update docs to include image feature extraction instructions * remove pooling layer from IJepaModel in image classification class * [run-slow] ijepa * remove pooling layer from IJepaModel constructor * update docs * [run-slow] ijepa * [run-slow] ijepa * small changes * [run-slow] ijepa * style adjustments * update copyright in init file * adjust modular ijepa * [run-slow] ijepa	2024-12-05 16:14:46 +01:00
Mohamed Mekkouri	95a855e212	Deprecate quanto and switch to optimum-quanto (#35001 ) * deprecate quanto * fix style	2024-12-05 16:11:09 +01:00
Isotr0py	482cb28a18	Fix `tie_word_embeddings` handling for GGUF models (#35085 ) * fix tie_word_embeddings Signed-off-by: Isotr0py <2037008807@qq.com> * fix Signed-off-by: Isotr0py <2037008807@qq.com> --------- Signed-off-by: Isotr0py <2037008807@qq.com>	2024-12-05 16:00:41 +01:00
Cyril Vallez	35447054f5	Update Mistral conversion script (#34829 ) * Update convert_mistral_weights_to_hf.py * Update convert_mistral_weights_to_hf.py * Update convert_mistral_weights_to_hf.py	2024-12-05 15:47:20 +01:00
Arthur	93f87d3cf5	[`tokenizers`] bump to 0.21 (#34972 ) bump to 0.21	2024-12-05 15:46:02 +01:00
eustlb	54aae121eb	[Whisper] Fix whisper tokenizer (#34537 ) * handle single timestamp ending * include last timestamp token * handle single timestamp ending * avoid floating points arithm limitations * ensure float64 operations * new test * make fixup * make copies * handle edge case double tokens ending with different tokens * handle single timestamp ending * make fixup * handle conditioning on prev segments * fix * Update src/transformers/models/whisper/generation_whisper.py Co-authored-by: Yoach Lacombe <52246514+ylacombe@users.noreply.github.com> * [run-slow] whisper * don't call item() to avoid unnecessary sync * fix --------- Co-authored-by: Yoach Lacombe <52246514+ylacombe@users.noreply.github.com> Co-authored-by: Eustache Le Bihan <eustlb@users.noreply.huggingface.co>	2024-12-05 13:46:29 +01:00
Yih-Dar	beb2c66ec3	Informative (#35059 ) * fix * fix * fix * fix --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2024-12-05 09:50:27 +01:00
Steven Liu	1ed1de2fec	[docs] Increase visibility of torch_dtype="auto" (#35067 ) * auto-dtype * feedback	2024-12-04 09:18:44 -08:00
Fanli Lin	baa3b22137	[docs] add a comment that offloading requires CUDA GPU (#35055 ) * add commen to offloading * Update docs/source/en/kv_cache.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2024-12-04 07:48:34 -08:00
Cyril Vallez	1da1e0d7f2	Support for easier multimodal use of modular (#35056 ) * update modular and add examples * style * improve example comments * style * fix small logic issue for imports * fix relative order issue when files do not make sense * Improve comments * trigger CIs	2024-12-04 15:13:11 +01:00
Anton Vlasjuk	46df859975	[`GPTNeoX`] Flex Attention + Refactor (#34896 ) * gpt neox flex attention + refactor * some formatting * small fix on dropout * add assertion on flex attn test * flaky ci :( * add head mask support * style * handle dtype, replace torch where * fixup flex with output attns * code review and several other fixes * Update src/transformers/modeling_utils.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * style * remove unnecessary comment * remove incorrect comment * make flex attn check more agnostic tor versions and centralized * change peft input dtype check to value since q and k could be affected by other stuff like RoPE * i forgor * flaky * code review and small fixes * Update src/transformers/models/gpt_neox/modeling_gpt_neox.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2024-12-04 14:48:28 +01:00
Vladislav Bronzov	accb7204f9	Add Pytorch Tensor Parallel support for Qwen2, Qwen2Moe, Starcoder2 (#35007 ) * add base tp plan for qwen2 and qwen2moe * add parallel tp for starcoder2 * fix modular conversion * add infer dim for qkv states * Update src/transformers/models/qwen2_moe/configuration_qwen2_moe.py --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2024-12-04 14:43:36 +01:00
Tianshu Wang	c7a109ec81	Fix `pad_token_tensor` is None in warning (#34005 ) Fix pad_token_tensor is None in warning	2024-12-04 11:15:25 +01:00
Fanli Lin	329f5dbf97	[docs] use device-agnostic API instead of hard-coded cuda (#35048 ) replace cuda	2024-12-03 10:54:15 -08:00
Fanli Lin	b8cdc262d5	[docs] use device-agnostic instead of `cuda` (#35047 ) * fix on xpu * [run_all] * add the missing import for Image lib * add more devices in comment * bug fix * replace cuda	2024-12-03 10:53:45 -08:00
wwwbai	346597b644	Translate community.md into Chinese (#35013 ) * community translation * Update docs/source/zh/community.md Co-authored-by: Isotr0py <2037008807@qq.com> --------- Co-authored-by: Isotr0py <2037008807@qq.com>	2024-12-03 10:22:02 -08:00
Fanli Lin	3deaa8179d	[docs] fix example code bug (#35054 ) fix code bug	2024-12-03 09:18:39 -08:00
Wang, Yi	125de41643	fix speecht5 failure issue in test_peft_gradient_checkpointing_enable… (#34454 ) * fix speecht5 failure issue in test_peft_gradient_checkpointing_enable_disable Signed-off-by: Wang, Yi <yi.a.wang@intel.com> * [run-slow] speecht5 --------- Signed-off-by: Wang, Yi <yi.a.wang@intel.com> Co-authored-by: Matt <rocketknight1@gmail.com>	2024-12-03 13:58:54 +00:00
Yih-Dar	7a7f27697a	Fix `BertGeneration` (#35043 ) fix Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2024-12-03 13:56:59 +01:00
Aymeric Roucher	901f504580	Add token cost + runtime monitoring to Agent and HfEngine children (#34548 ) * Add monitoring to Agent and HfEngine children	2024-12-03 13:14:52 +01:00
Cyril Vallez	ee37bf0d95	Automatic compilation in generate: do not rely on inner function (#34923 ) * compiled forward in PreTrainedModel * update * style * update name * trigger CIs * Add way to use custom compile args * style * switch parameterization to generation_config * Add to inits * Update configuration_utils.py * inits * style * docs * style * Update configuration_utils.py * back without dataclass for repo consistency * Update configuration_utils.py * style * style * style once again * add config serialization * update * true dataclass * trigger CIs * merge compile methods + remove serialization of compile config	2024-12-03 11:20:31 +01:00
wwwbai	f9c7e6021e	Translate bertlogy.md into Chinese (#34908 ) * bertology translation * Update docs/source/zh/_toctree.yml Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/zh/bertology.md Co-authored-by: blueingman <15329507600@163.com> * Update docs/source/zh/bertology.md Co-authored-by: blueingman <15329507600@163.com> * Update docs/source/zh/bertology.md Co-authored-by: Isotr0py <2037008807@qq.com> * Update docs/source/zh/bertology.md Co-authored-by: Isotr0py <2037008807@qq.com> --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> Co-authored-by: blueingman <15329507600@163.com> Co-authored-by: Isotr0py <2037008807@qq.com>	2024-12-02 11:42:40 -08:00
Fanli Lin	527dc04e46	[docs] add the missing import for Image and bug fix (#34776 ) * add the missing import for Image lib * add more devices in comment * bug fix	2024-12-02 11:40:20 -08:00
Ahmed Almaghz	4955e4e638	[i18n-ar] Translated file : `docs/source/ar/notebooks.md` into Arabic (#33049 ) * Add docs/source/ar/notebooks.md to Add_docs_source_ar_notebooks.md * Update notebooks.md * Update _toctree.yml	2024-12-02 11:40:04 -08:00
secrettoad	f0dec874f0	add docstring example for compute_loss_func (#35020 )	2024-12-02 11:39:09 -08:00
Henry Hyeonmok Ko	31299670cd	Multiple typo fixes in Tutorials docs (#35035 ) * Fixed typo in multi gpu docs and OLMoE version * Fixed typos in docs for agents, agents advanced, knowledge distillation, and image feature extraction * Fixed incorrect usage of model.image_guided_detection in zero shot object detection docs	2024-12-02 15:26:34 +00:00
Dmitry Rogozhkin	31830474bf	Fix `test_eager_matches_sdpa_inference` for `XPU` backend (#34889 ) * Use torch.nn.attention.sdpa_kernel instead of deprecated torch.backends.cuda.sdp_kernel Signed-off-by: Dmitry Rogozhkin <dmitry.v.rogozhkin@intel.com> * Fix test_eager_matches_sdpa_inference for XPU backend As of PyTorch 2.5 XPU backend supports only torch.nn.attention.SDPBackend.MATH which is implemented on PyTorch level using aten operators and is device agnostic with respect to implementation of each aten operator. Thus, we can reuse CUDA (or CPU) MATH weights for XPU. Fixes: #34888 Signed-off-by: Dmitry Rogozhkin <dmitry.v.rogozhkin@intel.com> * Use torch.amp.autocast instead of deprecated torch.cuda.amp.autocast in nemotron Signed-off-by: Dmitry Rogozhkin <dmitry.v.rogozhkin@intel.com> --------- Signed-off-by: Dmitry Rogozhkin <dmitry.v.rogozhkin@intel.com>	2024-12-02 16:21:04 +01:00
Jacky Lee	f41d5d8f74	Add type hints for forward functions in Gemma2 (#35034 ) * feat: add gemma2 type hints * fix: mask is optional	2024-12-02 14:03:36 +00:00
Bojun Feng	7b5f76e32e	Typo in warning switching to optimum-quanto (#35028 ) fix typos	2024-12-02 13:47:05 +00:00
milesial	c24c79ebf9	Optimize memory usage of mllama encoder (#34930 ) mllama encoder memory optimization	2024-12-02 11:46:45 +01:00
Weize Chen	9ab8c5b503	fix variable undefined bug when return_tensors is not specified in llava processing (#34953 ) * fix variable undefined bug when return_tensors is not specified in llava processor * improve readability	2024-12-02 11:44:42 +01:00
Joshua Lochner	3480cbb97e	Only cast `cu_seqlens` when tracing (#35016 ) * Only cast `cu_seqlens` when tracing * Formatting	2024-12-02 11:39:39 +01:00
Alvaro Bartolome	19dabe9636	Update `FillMaskPipeline.__call__` signature and docstring (#35006 ) Update `FillMaskPipeline.__call__` - Remove unused `*args` - Update docstring with `inputs` over `args`	2024-11-29 13:44:56 +00:00
Samuel Larkin	f7427f58ed	fix: double verbs (#35008 )	2024-11-29 13:19:57 +00:00

... 36 37 38 39 40 ...

19383 Commits