transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-18 20:18:24 +06:00

Author	SHA1	Message	Date
Zach Mueller	863e2562d8	Make clearer about zero_init requirements (#29879 ) * Docstring to note about zero init * Check for accelerate * Change conditional return * Tweak * Add new accelerate-specific zero3 check * Fix import * Revert to RTFM * Update src/transformers/modeling_utils.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> --------- Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2024-04-03 13:37:52 -04:00
Arthur	695d823323	[`Main CIs`] Fix the red cis (#30022 ) * fix * sort imports	2024-04-03 19:34:39 +02:00
Raushan Turganbay	c10b5dd25e	Superpoint imports fix (#29898 ) quick fix	2024-04-03 18:32:01 +01:00
Steven Liu	34bfe95af5	[docs] Fix audio file (#30006 ) new audio file	2024-04-03 10:05:15 -07:00
Raushan Turganbay	cc75f1ac73	Fix vipllava for generation (#29874 ) * fix vipllava generation * consistent llava code * revert llava tests changes	2024-04-03 17:00:08 +01:00
Ondřej Cífka	240e10626b	Fix probability computation in `WhisperNoSpeechDetection` when recomputing scores (#29248 ) * Fix is_scores_logprobs in WhisperNoSpeechDetection * Add test_whisper_longform_no_speech_detection * Fix typo	2024-04-03 17:53:07 +02:00
Ondřej Cífka	bcd42c4af9	Fix `kwargs` handling in `generate_with_fallback` (#29225 ) * Fix generate_with_fallback *kwargs Change pop to get * Delete keys from kwargs to prevent overriding generation_config * Revert to passing kwargs by reference, but make a (shallow) copy * dict -> copy.copy * Add test_whisper_longform_multi_batch_beam	2024-04-03 17:51:03 +02:00
Ren Xuancheng	851f253f4d	Fix Qwen2Tokenizer (#29929 ) qwen2: fixed tokens starting with # in slow tokenizer; add tests Co-authored-by: jklj077 <17811943+jklj077@users.noreply.github.com>	2024-04-03 17:42:43 +02:00
Miguel Almeida	17b06e2c66	Fix Swinv2ForImageClassification NaN output (#29981 ) To address the issue of NaN logit outputs for certain combinations of the `image_size`, `patch_size` and `depths` configuration parameters, an assertion was made to ensure that the resulting `window_size` field in the model's Self Attention class is greater than 1, preventing divisions by zero in the normalization of `relative_coords_table`. Fix: #28675	2024-04-03 14:54:45 +01:00
fxmarty	81642d2b51	Make EncodecModel.decode ONNX exportable (#29913 ) * fix encodec onnx export for musicgen * simplification * fix quality * better style	2024-04-03 17:11:01 +08:00
Yih-Dar	b44df05bc0	Update `tests/utils/tiny_model_summary.json` (#29941 ) update Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2024-04-03 09:25:01 +02:00
Mario Šaško	fce52cefa7	Fix `remove_columns` in `text-classification` example (#29351 )	2024-04-02 19:15:27 +02:00
Joao Gante	5080ab12c8	Generate: fix logits processors doctests (#29718 ) * fix norm * fix logits processors doctests	2024-04-02 17:18:31 +01:00
Nicolas Patry	9b0a8ea7d1	Hard error when ignoring tensors. (#27484 ) (#29906 ) * Hard error when ignoring tensors. (#27484) * [WIP] Hard error when ignoring tensors. * Better selection/error when saving a checkpoint. - Find all names we should normally drop (those are in the transformers config) - Find all disjoint tensors (for those we can safely trigger a copy to get rid of the sharing before saving) - Clone those disjoint tensors getting rid of the issue - Find all identical names (those should be declared in the config but we try to find them all anyway.) - For all identical names: - If they are in the config, just ignore them everything is fine - If they are not, warn about them. - For all remainder tensors which are shared yet neither identical NOR disjoint. raise a hard error. * Adding a failing test on `main` that passes here. * We don't need to keep the subfolder logic in this test. * Apply suggestions from code review Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Add small tests. * Dead variable. * Fixup. * Fixing tied_Weights_keys on generic models. * Fixup + T5 encoder/decoder tying (with different layers) * Code quality. * Dynamic member. * trigger * Fixing encoder name for other types of encoder/decoder combos. * Fix scoping. * Update .github/workflows/self-scheduled.yml Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Fixing the tied_weights after the call. --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2024-04-02 16:59:05 +02:00
Minsub Lee (Matt)	15cd68713d	Fix `skip_special_tokens` for `Wav2Vec2CTCTokenizer._decode` (#29311 ) * Fix skip_special_tokens process for Wav2Vec2CTCTokenizer._decode * Fix skip_special_tokens for Wav2Vec2CTCTokenizer._decode * Exclude pad_token filtering since it is used as CTC-blank token * Add small test for skip_special_tokens * Update decoding test for added new token	2024-04-02 16:55:11 +02:00
Michael	cb5927ca8f	[Docs] Make an ordered list prettier in add_tensorflow_model.md (#29949 )	2024-04-02 12:37:56 +01:00
Yoach Lacombe	0d04b1e25a	Add Flash Attention 2 support to Musicgen and Musicgen Melody (#29939 ) * add FA2 to o.g Musicgen * make style * add FA2 support to Musicgen Melody * add generation FA2 tests to o.g Musicgen * make style and fix copies * add Musicgen to FA2 docs + deprecate list * add sdpa supports to Musicgen's * make style and fix copies * refactor attention implementation arguments * add Copied from to sdpa tests * add copied form in sdpa tests melody * add copied for FA2 generation tests * add FA2 inference copied from * make style	2024-04-02 11:23:49 +01:00
théo gigant	fed27ffc7e	Adding FlaxNoRepeatNGramLogitsProcessor (#29677 ) * fix issue with logit processor in beam search in Flax * adding FlaxNoRepeatNGramLogitsProcessor class + unit test * style correction and code verification * add FlaxNoRepeatNGramLogitsProcessor to the test_processor_list and test_processor_list_jitted tests * fix an issue where ngrams are banned only if they appear ==1 time + update description of get_previous_ngrams * replace non-jit compatible masking of ngrams that are not yet generated with jittable version * Revert "fix issue with logit processor in beam search in Flax" This reverts commit `09b70d7e4d`. * add FlaxNoRepeatNGramLogitsProcessor to _get_logits_processor * change the method of casting to boolean of banned tokens indices * fix code style * remove some useless operations + significantly faster computation of update indices using jax.lax.fori_loop * remove useless loop iterations * set some variables that were calculated and used multiple times * fix format	2024-04-02 11:39:33 +02:00
Marc Sun	33288ff150	[bnb] Fix bug in `_replace_with_bnb_linear` (#29958 ) fix bug	2024-04-02 11:18:03 +02:00
Hovnatan Karapetyan	416711c3ea	Fix 29807 sinusoidal positional encodings in Flaubert, Informer and XLM (#29904 ) * Fix sinusoidal_embeddings in FlaubertModel * Fix for Informer * Fix for XLM * Move sinusoidal emb for XLM * Move sinusoidal emb for Flaubert * Small cleanup * Add comments on tests code copied from * Add with Distilbert->	2024-04-02 10:27:26 +02:00
Arthur	83b26dd79d	[`generate`] fix breaking change for patch (#29976 ) * fix bug and add tests * nit * otherway to get the cur len instead of attention mask * more places where this might have been broken * nit * oups * inputs_embeds vs input_embeds * test generated outptus * style * nit * fix * skip failing biogpt	2024-04-02 09:51:45 +02:00
Steven Liu	096f304695	[docs] Big model loading (#29920 ) * update * feedback	2024-04-01 18:47:32 -07:00
Joao Gante	c9f6e5e351	Generate: move misplaced test (#29902 )	2024-04-01 12:45:25 +01:00
Fanli Lin	e4f5b57a3b	[tests] fix the wrong output in `ImageToTextPipelineTests.test_conditional_generation_llava` (#29975 ) bug fix	2024-04-01 13:08:39 +02:00
Arthur	fa2c49b00b	Fix copies main ci (#29979 ) * fix copies * nit * style * Update utils/check_copies.py	2024-04-01 12:43:58 +02:00
Yoach Lacombe	569f6c7d43	Fix FA2 tests (#29909 ) * fix FA2 tests * refactor inference test name	2024-04-01 07:51:00 +00:00
Zach Mueller	3b8e2932ce	Rework tests to compare trainer checkpoint args (#29883 ) * Start rework * Fix failing test * Include max * Update src/transformers/trainer.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2024-03-30 22:19:17 -04:00
TechxGenus	6e584070d4	[`BC`] Fix BC for AWQ quant (#29965 ) fix awq quant	2024-03-30 19:37:25 +01:00
Bo Zheng	46d636818b	Update model card and link of blog post. (#29928 ) * Update qwen2_moe.md * update link of blogpost. * fixup --------- Co-authored-by: bozheng-hit <dsoul0621@gmail.com>	2024-03-30 17:49:03 +01:00
Gary Wang	f6701bc664	Reset alarm signal when the function is ended (#29706 ) Fixes #29690	2024-03-30 17:41:27 +01:00
Alexander Jipa	e644b60038	fix: get mlflow version from mlflow-skinny (#29918 ) Co-authored-by: Alexander Jipa <azzhipa@amazon.com>	2024-03-30 17:38:29 +01:00
Jacky Lee	156d30da94	Add warning message for `run_qa.py` (#29867 ) * improve: error message for best model metric * update: raise warning instead of error	2024-03-30 17:02:31 +01:00
Jacky Lee	6fd93fe93a	Fix rope theta for OpenLlama (#29893 ) fix: rope_theta for open llama	2024-03-30 16:30:52 +01:00
fzyzcjy	5ad7f17002	Super tiny fix 12 typos about "with with" (#29926 ) * with with * style	2024-03-29 14:31:31 +00:00
Yih-Dar	43d17c1836	Mark `test_eager_matches_sdpa_generate` flaky for some models (#29479 ) * fix * revert for qwen2 * revert for qwen2 * update * update --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2024-03-29 11:51:20 +01:00
MariaHei	ba56ed0869	Update installs in image classification doc (#29947 ) Trainer with PyTorch now requires accelerate to be installed. Partly resolves huggingface/transformers#29174	2024-03-28 14:26:27 -07:00
Arthur	536ea2aca2	[`LlamaSlowConverter`] Slow to Fast better support (#29797 ) * fix * fix test * style * nit * rather rely on concert token to id * fix quality * Update src/transformers/convert_slow_tokenizer.py	2024-03-28 16:19:32 +01:00
VINAYAKK GARG	e203646871	Fix doc issue #29758 in DebertaV2Config class (#29842 ) Fix doc issue in DebertaV2Config class Co-authored-by: Vinayakk Garg <vigar@akamai.com>	2024-03-28 14:49:57 +00:00
Arthur	2bbbf1be5b	[`BC`] Fix BC for other libraries (#29934 ) * fi xbc? * nit	2024-03-28 15:13:23 +01:00
Yu Chin Fabian Lim	4df5b9b4b2	Allow GradientAccumulationPlugin to be configured from AcceleratorConfig (#29589 ) * add gradient_accumulation_kwargs to AcceleratorConfig * add suggestions from @muellerzr to docstrings, new behavior and tests * Documentation suggestions from @muellerz Co-authored-by: Zach Mueller <muellerzr@gmail.com> * addressed @muellerzr comments regarding tests and test utils * moved accelerate version to top of file. * @muellerzr's variable fix Co-authored-by: Zach Mueller <muellerzr@gmail.com> * address @amyeroberts. fix tests and docstrings * address @amyeroberts additional suggestions --------- Co-authored-by: Yu Chin Fabian Lim <flim@sg.ibm.com> Co-authored-by: Zach Mueller <muellerzr@gmail.com>	2024-03-28 14:01:40 +00:00
Arthur	a2a7f71604	[ `TokenizationLlama`] fix the way we convert tokens to strings to keep leading spaces 🚨 breaking fix (#29453 ) * nit * update test and fix test * fixup	2024-03-28 13:58:40 +01:00
Arthur	e677479c81	[`Mamba`] from pretrained issue with `self.embeddings` (#29851 ) * nit * update * oups * Update src/transformers/models/mamba/modeling_mamba.py Co-authored-by: Lysandre Debut <hi@lysand.re> --------- Co-authored-by: Lysandre Debut <hi@lysand.re>	2024-03-28 13:54:51 +01:00
Joao Gante	441de62f49	RoPE models: add numerical sanity-check test for RoPE scaling (#29808 ) * add hard rope scaling test * make fixup * quick rope scaling tests * add copy statements	2024-03-28 11:25:50 +00:00
Christopher Keibel	aac7099c92	add functions to inspect model and optimizer status to trainer.py (#29838 ) * add functions to get number of params which require grad, get optimizer group for parameters and get learning rates of param groups to trainer.py * add tests and raise ValueError when optimizer is None * add second layer to test and freeze its weigths * check if torch is available before running tests * use decorator to check if torch is available Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * fix test indentation Co-authored-by: Zach Mueller <muellerzr@gmail.com> --------- Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> Co-authored-by: Zach Mueller <muellerzr@gmail.com>	2024-03-28 10:37:16 +00:00
amyeroberts	855b95ce34	Safe import of LRScheduler (#29919 ) * Safe import of LRScheduler * Update src/transformers/trainer_pt_utils.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update src/transformers/trainer_pt_utils.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Fix up --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2024-03-28 09:54:51 +00:00
Aymeric Roucher	c9d2e855ea	Add beam search visualizer to the doc (#29876 )	2024-03-28 09:54:08 +00:00
Joao Gante	248d5d23a2	Tests: replace `torch.testing.assert_allclose` by `torch.testing.assert_close` (#29915 ) * replace torch.testing.assert_allclose by torch.testing.assert_close * missing atol rtol	2024-03-28 09:53:31 +00:00
Fanli Lin	7c19fafe44	[doc] fix some typos and add `xpu` to the testing documentation (#29894 ) fix typo	2024-03-28 09:42:49 +00:00
Eduardo Pacheco	22d159ddf9	Adding Flash Attention 2 Support for GPT2 (#29226 ) * First commit to add flash attention 2 for GPT-2 * more improvements * Make GPT2 pass tests and fixed Decison Transformers copies * Fixed missing arg * fix copies * Added expected speedup * Update src/transformers/models/gpt2/modeling_gpt2.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update src/transformers/models/gpt2/modeling_gpt2.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update src/transformers/models/gpt2/modeling_gpt2.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Added test * Fixed attn attribute * Update docs/source/en/model_doc/gpt2.md Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update docs/source/en/model_doc/gpt2.md Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update Decision transformer attentions * More updates * Passing tests * Fix copies * Fix copies part 2 * Decision transformer updates * Update src/transformers/models/gpt2/modeling_gpt2.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Fix copies * Decision transformer not supporting flash attn * Addressed comments * Addressed comments * Addressed comments --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2024-03-28 09:31:24 +00:00
Arthur	3a7e68362b	[`pipeline`]. Zero shot add doc warning (#29845 ) * add doc warning * fix build pr	2024-03-28 09:10:26 +01:00

... 11 12 13 14 15 ...

16108 Commits