transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-08-03 03:31:05 +06:00

Author	SHA1	Message	Date
Konstantin Dobler	650a71e157	Support ratios for `logging_steps`, `eval_steps`, and `save_steps` (#23235 ) * Ratio option for `logging_steps`, `eval_steps`, `save_steps` * Add guards if arguments are not set * Add more detailed comments + formatting * Update src/transformers/training_args.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/training_args.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/training_args.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Convert args values to `int` if bigger than 1 * `black` * `make fixup` --------- Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2023-05-09 13:05:13 -04:00
Nicolas Patry	c34a525d2f	Proposed fix for TF example now running on safetensors. (#23208 ) * Proposed fix for TF example now running on safetensors. * Adding more warnings and returning keys. * Trigger CI * Trigger CI --------- Co-authored-by: Sylvain Gugger <Sylvain.gugger@gmail.com>	2023-05-09 13:04:27 -04:00
Sylvain Gugger	b4d4d6fe87	Add RWKV-4 (#22797 ) * First draft of RWKV-4 * Add support for generate * Style post-rebase * Properly use state * Write doc * Fix doc * More math * Add model to README, dummies and clean config * Fix init * multiple fixes: - fix common tests - fix configuraion default values - add CI test for checking state computation - fix some CI tests * correct tokenizer * some tweaks - fix config docstring - fix failing tests * fix CI tests - add output_attention / output_hidden_states - override test_initialization - fix failing CIs * fix conversion script - fix sharded case - add new arguments * add slow tests + more fixes on conversion script * add another test * final fixes * change single name variable * add mock attention mask for pipeline to work * correct eos token id * fix nits * add checkpoints * Apply suggestions from code review Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * add `tie_word_embeddings` in docstring * change tensor name * fix final nits * Trigger CI --------- Co-authored-by: younesbelkada <younesbelkada@gmail.com> Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com> Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2023-05-09 13:04:10 -04:00
Rustin Welter	9a50cb6195	Add Japanese translation to accelerate.mdx (#23232 ) Co-authored-by: rustinwelter <rustinwelter.alwp9@slmails.com>	2023-05-09 10:51:43 -04:00
Sebastian	1a8f61110e	fix: Update run_qa.py to work with deepset/germanquad (#23225 ) Call str on id to make sure any ints are converted into the expected format for squad datasets	2023-05-09 09:20:10 -04:00
Furkan Akkurt	51ae566511	Fix typo ; Update output.mdx (#23227 )	2023-05-09 09:19:38 -04:00
dumpmemory	e02a8065e0	make opt checkpoint dir name correct (#21660 ) make opt checkpoint dir name corrent following `100b522bb8/megatron/checkpointing.py (L117)`	2023-05-09 09:14:02 -04:00
Matthijs Hollemans	7f91950901	audio_utils improvements (#21998 ) * silly change to allow making a PR * clean up doc comments * simplify hertz_to_mel and mel_to_hertz * fixup * clean up power_to_db * also add amplitude_to_db * move functions * clean up mel_filter_bank * fixup * credit librosa & torchaudio authors * add unit tests * tests for power_to_db and amplitude_to_db * add mel_filter_bank tests * rewrite STFT * add convenience spectrogram function * missing transpose * fewer transposes * add integration test to M-CTC-T * frame length can be either window or FFT length * rewrite stft API * add preemphasis coefficient * move argument * add log option to spectrogram * replace M-CTC-T feature extractor * fix api thing * replace whisper STFT * replace whisper mel filters * replace tvlt's stft * allow alternate window names * replace speecht5 stft * fixup * fix integration tests * fix doc comments * remove manual FFT length calculation * fix docs * go away, deprecation warnings * combine everything into spectrogram function * add deprecated functions back * fixup	2023-05-09 09:10:17 -04:00
NielsRogge	431b04d8c4	[SAM] Add resources (#23224 ) Add resources	2023-05-09 08:58:19 -04:00
Sylvain Gugger	006da469dd	Pin tensorflow-probability (#23220 ) * Pin tensorflow-probability * [all-test] * [all-test] Fix syntax for bash	2023-05-08 18:36:22 -04:00
Connor Henderson	188a8bfccc	docs: Fix broken link in 'How to add a model...' (#23216 ) fix link	2023-05-08 14:56:42 -04:00
Sylvain Gugger	94056b57be	New version of Accelerate for the Trainer (#23204 )	2023-05-08 09:47:08 -04:00
Sylvain Gugger	fd6970bc56	Skip failing test	2023-05-08 08:52:44 -04:00
Orr Zohar	843fdf2e42	Fixing class embedding selection in owl-vit (#23157 ) fixing class embedding selection in owl-vit	2023-05-08 07:35:04 -04:00
Joao Gante	bbfb9fc22b	Generate: starcoder 🤜 🤛 assisted generation (#23182 ) * starcoder has joined the chat * indexing that works for all	2023-05-08 10:45:40 +01:00
Robert Baruch	dbc12269ed	Fix hf_argparser.parse_json_file to open file with utf-8 encoding, close file when finished (#23194 ) * Open json args in utf-8 encoding, close file when finished * black formatted	2023-05-07 19:06:24 -04:00
Bartosz Szmelczynski	6f8a02844a	fix random attention for pytorch's bigbird/pegasus_bigbird (#23056 ) * fix random attention usage for bigbird and pegasus_bigbird * remove staticmethod, update tests target valus * revert style changes	2023-05-07 18:55:04 -04:00
Ashwin Mathur	ef0c380c12	Update LLaMA docs with arxiv link (#23191 ) * Update docs with arxiv link * Update llama model docs	2023-05-07 18:52:44 -04:00
cyy	ef42c2c487	search buffers for dtype (#23159 )	2023-05-06 11:41:08 -04:00
raghavanone	312b104ff6	Add FlaxWhisperForAudioClassification model (#23173 ) * Add FlaxWhisperForAudioClassification model * Add models to init * Add models to init * Fix copies * Fix automapping * Fix failing test	2023-05-05 13:23:46 -04:00
Ashwin Mathur	fc6c8b0eaa	Add `no_trainer` scripts to pre-train Vision Transformers (#23156 ) * Add run_mim_no_trainer.py draft from #20412 Add parse_args method and copy over other dependencies Add Method call for sending telemetry Initialize Accelerator Make one log on every process Set seed and Handle repository creation Initialize dataset and Set validation split Create Config Adapt Config Update Config Create Feature Extractor Create model Set column names Create transforms Create mask generator Create method to preprocess images Shuffle datasets if needed and set transforms Create Dataloaders Add optimizer Add learning rate scheduler Prepare everything with our accelerator Tie weights for TPU training Recalculate training steps and training epochs Set accelerator checkpointing steps Initialize trackers and store configuration Set total batch size Fix typo: mlm -> mim Log info at the start of training Load in the weights and states from previous save update the progress_bar if load from checkpoint Define train loop Add evaluation loop to training Add to parse_args method Push repo to hub Save accelerator state End training and save model and feature extractor Remove unused imports Fix trailing whitespace * Update code based on comments, Rename feature_extractor to image_processor * Fix linting * Add argument for learning rate * Add argument for setting number of training epochs * Remove incorrect logger argument * Convert max_train_steps to int for tqdm --------- Co-authored-by: Saad Mahmud <shuvro.mahmud79@gmail.com>	2023-05-05 13:22:49 -04:00
Connor Henderson	17083b9b84	fix: Passing language as acronym to Whisper generate (#23141 ) * add fix * address comments * remove error formatting	2023-05-05 11:52:19 -04:00
Gabriel Yang	40082d598b	🌐 [i18n-KO] docs: ko: Translate `multiple_choice.mdx` (#23064 ) * update doctree * doc: ko: translate multiple choice * Update reviews	2023-05-05 11:36:56 -04:00
Andrei Filatov	77412343c8	fixed whisper positional encoding (#23167 )	2023-05-05 11:36:15 -04:00
Perry Huang	1b9c352e55	Add TrOCR resources (#23142 ) * Add TrOCR resources * Made fixes suggested by stevhliu	2023-05-05 11:29:20 -04:00
Sylvain Gugger	01734dba84	Revert "Add FlaxWhisperForAudioClassification model" (#23154 ) Revert "Add FlaxWhisperForAudioClassification model (#22883)" This reverts commit `c8f2c5c56e`.	2023-05-04 13:47:07 -04:00
Joao Gante	b369e507aa	Generate: text generation pipeline no longer emits `max_length` warning when it is not set (#23139 )	2023-05-04 18:36:23 +01:00
Maria Khalusova	516dc6305f	[docs] Text to speech task guide (#23107 ) * First draft * Some polishing * Text polishing * added TOC entry for TTS * make style * added links to images * fixed links to images * Apply suggestions from code review Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * feedback addressed * feedback from Matthijs addresed * Update docs/source/en/tasks/text-to-speech.mdx Co-authored-by: Matthijs Hollemans <mail@hollance.com> --------- Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: Matthijs Hollemans <mail@hollance.com>	2023-05-04 13:17:13 -04:00
raghavanone	c8f2c5c56e	Add FlaxWhisperForAudioClassification model (#22883 ) * Add FlaxWhisperForAudioClassification model * Add models to init * Add models to init * Fix copies * Fix automapping	2023-05-04 13:00:16 -04:00
Sylvain Gugger	3341bb41cd	Pin urllib3	2023-05-04 12:00:22 -04:00
Younes Belkada	57ffd8ab4c	[`GPT-J`] Fix causal mask dtype (#23147 ) * fix #23136 * better fix * same fix for `masked_bias`	2023-05-04 16:31:19 +02:00
peter-sk	83b38fbea8	GPTNeoXForQuestionAnswering (#23059 ) * first draft - gives index error in question_answering.py * maturing * no labels * pipeline should know about QA * fixing checks * formatting * fixed docstring * initial commit * formatting * adding the class to many places * towards less unhappy checks * nearly there * and gpt neox for qa * use right model * forgot this one * base_model_prefix is "gpt_neox" for GPTNeoX* models * unnecessary stuff * Update src/transformers/models/gpt_neox/modeling_gpt_neox.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * format * Update src/transformers/models/gpt_neox/modeling_gpt_neox.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * removed gpt2 stuff --------- Co-authored-by: Prof. Peter Schneider-Kamp <jps@ordbogen.com> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2023-05-04 10:15:15 -04:00
peter-sk	510ad0a8b8	gpt2 multi-gpu fix (#23149 ) Co-authored-by: Prof. Peter Schneider-Kamp <jps@ordbogen.com>	2023-05-04 09:58:38 -04:00
Qingyang Wu	adb0760b5f	fix resume fsdp (#23111 ) * fix resume fsdp * fix rank 0 loading * fix style and quality	2023-05-04 09:57:32 -04:00
Victor Geislinger	3b74889e8f	Remove typo in perf_train_gpu_many.mdx (#23144 ) - Excess `w` in the word `bottom`	2023-05-04 09:56:45 -04:00
digger-yu	5eeb556484	fix spelling error (#23143 ) change referrred to referred	2023-05-04 09:56:28 -04:00
amyeroberts	90e8263d91	Add methods to update and verify out_features out_indices (#23031 ) * Add methods to update and verify out_features out_indices * Safe update for config attributes * Fix function names * Save config correctly * PR comments - use property setters * PR comment - directly set attributes * Update test * Add updates to recently merged focalnet backbone	2023-05-04 10:15:06 +01:00
peter-sk	78b7debf56	GPTNeoForQuestionAnswering (#23057 ) * first draft - gives index error in question_answering.py * maturing * no labels * pipeline should know about QA * fixing checks * formatting * fixed docstring * initial commit * formatting * adding the class to many places * towards less unhappy checks * nearly there * Update src/transformers/models/gpt_neo/modeling_gpt_neo.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * avoid error * moving to device of star/end_logits --------- Co-authored-by: Prof. Peter Schneider-Kamp <jps@ordbogen.com> Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2023-05-03 15:59:19 -04:00
Robert Stone	b6933d76d2	Tidy Pytorch GLUE benchmark example (#23134 ) Migration to Evaluate for metric is not quite complete	2023-05-03 15:50:41 -04:00
Alara Dirik	b0a78091a5	Remove redundant print statements (#23133 ) remove redundant print statements	2023-05-03 18:04:48 +01:00
regisss	e3ee45aa54	Enable to use custom tracer in FX `symbolic_trace` (#23105 ) * Enable to use custom tracer in FX `symbolic_trace` * Integrate feedback from review * Formatting Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> --------- Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2023-05-03 12:47:36 -04:00
Alara Dirik	441658dd6c	Add focalnet backbone (#23104 ) Adds FocalNet backbone to return features from all stages	2023-05-03 19:32:42 +03:00
Julien Chaumond	ca7eb27ed5	[doc] Try a few ≠ ways of linking to Papers, users, and org profiles (#22611 ) * [doc] Try a few ≠ ways of linking to Papers, users, and org profiles * Empty commit * Empty commit now that the backend is fixed --------- Co-authored-by: Lysandre <lysandre@huggingface.co>	2023-05-03 18:23:09 +02:00
Nayeon Han	fbe0178f08	docs: ko: update `_toctree.yml` (#23112 ) * docs: ko: update `_toctree.yml` * fix: ko: update toc * fix: resolve suggestions * fix: resolve build issue --------- Co-authored-by: Wonhyeong Seo <wonhseo@kakao.com>	2023-05-03 11:04:58 -04:00
Mayank Agarwal	c4e32e206f	Add support for beam search's num_return_sequencs flag in flax (#23082 ) * add code for numReturnSeq * add flax support for num return sequences * Make Fix up for changes * add test for num return sequences * lint	2023-05-03 10:50:34 -04:00
Xuehai Pan	ee4bc07474	Support union types `X \| Y` syntax for `HfArgumentParser` for Python 3.10+ (#23126 ) * Support union types `X \| Y` syntax for `HfArgumentParser` for Python 3.10+ * Add tests for PEP 604 for `HfArgumentParser` * Reorganize tests	2023-05-03 10:49:54 -04:00
Alara Dirik	56b8d49ddf	Fix ConvNext V2 paramater naming issue (#23122 ) Fixes the parameter naming issue in ConvNextV2GRN module	2023-05-03 17:21:27 +03:00
Samin Yasar	b53004fdce	Add resources for LayoutLmV2 and reformat documentation resources (#23115 ) * add resources for layoutlmv2 * remove 🌎 from some resources	2023-05-03 09:53:00 -04:00
Joao Gante	3a08dc63fd	Generate: better warnings with pipelines (#23128 )	2023-05-03 14:43:17 +01:00
Manuel	2a16d8b275	improve unclear documentation (#23123 )	2023-05-03 09:36:30 -04:00

1 2 3 4 5 ...

12807 Commits