transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-24 06:48:58 +06:00

Author	SHA1	Message	Date
Sylvain Gugger	87dd1a00ef	Fix metric computation in `run_glue_no_trainer` (#11569 )	2021-05-03 11:42:55 -04:00
Muktan	a721a5eefd	[Wav2vec2] Fixed tokenization mistakes while adding single-char tokens to tokenizer (#11538 ) * Fixed tokenization mistakes while adding single-char tokens to tokenizer * Added tests and Removed unnecessary comments. * finalize wav2vec2 tok * add more aggressive tests * Apply suggestions from code review * fix useless import Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>	2021-05-03 17:19:12 +02:00
NielsRogge	f3cf8ae7b3	Add LUKE (#11223 ) * Rebase with master * Minor bug fix in docs * Copy files from adding_luke_v2 and improve docs * change the default value of use_entity_aware_attention to True * remove word_hidden_states * fix head models * fix tests * fix the conversion script * add integration tests for the pretrained large model * improve docstring * Improve docs, make style * fix _init_weights for pytorch 1.8 * improve docs * fix tokenizer to construct entity sequence with [MASK] entity when entities=None * Make fix-copies * Make style & quality * Bug fixes * Add LukeTokenizer to init * Address most comments by @patil-suraj and @LysandreJik * rename _compute_extended_attention_mask to get_extended_attention_mask * add comments to LukeSelfAttention * fix the documentation of the tokenizer * address comments by @patil-suraj, @LysandreJik, and @sgugger * improve docs * Make style, quality and fix-copies * Improve docs * fix docs * add "entity_span_classification" task * update example code for LukeForEntitySpanClassification * improve docs * improve docs * improve the code example in luke.rst * rename the classification layer in LukeForEntityClassification from typing to classifier * add bias to the classifier in LukeForEntitySpanClassification * update docs to use fine-tuned hub models in code examples of the head models * update the example sentences * Make style & quality * Add require_torch to tokenizer tests * Add require_torch to tokenizer tests * Address comments by @sgugger and add community notebooks * Make fix-copies Co-authored-by: Ikuya Yamada <ikuya@ikuya.net>	2021-05-03 09:07:29 -04:00
Frederik Bode	6a11e4c2ad	fix the mlm longformer example by changing [MASK] to <mask> (#11559 )	2021-05-03 12:43:30 +01:00
Lysandre Debut	1c86157d9d	Remove `datasets` submodule. (#11563 )	2021-05-03 06:02:33 -04:00
Patrick von Platen	c448c01f25	[Wav2Vec2] Fix convert (#11562 ) * push * small change * correct other typo	2021-05-03 11:53:30 +02:00
Suraj Patil	623281aa12	[Flax BERT/Roberta] few small fixes (#11558 ) * small fixes * style	2021-05-03 10:35:06 +02:00
lewtun	a5d2967bd8	Fix examples in M2M100 docstrings (#11540 ) Replaces `tok` with `tokenizer` so examples can run with copy-paste	2021-05-03 10:56:31 +05:30
jingyihe	980208650a	Fixed docs for the shape of `scores` in `generate()` (#10057 ) * Fixed the doc for the shape of return scores tuples in generation_utils.py. * Fix the output shape of `scores` for `DecoderOnlyOutput`. * style fix	2021-05-02 10:10:47 +02:00
Stas Bekman	4e7bf94e72	[DeepSpeed] fp32 support (#11499 ) * prep for deepspeed==0.3.16 * new version * too soon * support and test fp32 mode * troubleshooting doc start * workaround no longer needed * add fp32 doc * style * cleanup, add tf32 note * clarify * release was made	2021-04-30 12:51:48 -07:00
Stas Bekman	282f3ac3ef	[debug utils] activation/weights underflow/overflow detector (#11274 ) * sync * add activation overflow debug utility * cleanup * document detect_overflow * import torch * add deprecation warning * Apply suggestions from code review Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * convert to rst, add note * add class * fix docs * improve the doc * rework to dump a lot more info about each frame * complete expansion * cleanup * format * cleanup * doesn't have to be transformers * Apply suggestions from code review Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * wrap long line * style Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2021-04-30 11:15:46 -07:00
Hamel Husain	804c2974d5	Improve task summary docs (#11513 ) * fix task summary docs * refactor to use model.config.id2label instead of list * fix nit * Update docs/source/task_summary.rst Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2021-04-30 09:06:47 -04:00
Sylvain Gugger	bc80f8bc37	Add Stas and Suraj as authors (#11526 )	2021-04-30 09:03:13 -04:00
Bhadresh Savani	84326a28f8	[Examples] Added support for test-file in QA examples with no trainer (#11510 ) * added support for test-file * fixed typo * added suggested changes * reformatted code * modifed files * fix post processing error * Trigger CI * removed extra lines	2021-04-30 09:02:50 -04:00
Lysandre Debut	af0692a2ca	Run model templates on master (#11527 )	2021-04-30 08:47:12 -04:00
Suraj Patil	57c8e822f7	reszie token embeds (#11524 )	2021-04-30 08:47:01 -04:00
Matt	20d6931e32	Update TF text classification example (#11496 ) Big refactor, fixes and multi-GPU/TPU support	2021-04-30 13:45:33 +01:00
bonniehyeon	8b945ef03e	Fix do_eval default value in training_args.py (#11511 ) * Fix do_eval default value in training_args.py * Update PULL_REQUEST_TEMPLATE.md	2021-04-30 08:35:12 -04:00
Takuya Makino	c2cd02ac62	Accepts BatchEncoding in LengthSampler (#11431 )	2021-04-30 08:27:46 -04:00
Shubham Sanghavi	30ede8994e	Implement Fast Tokenization for Deberta (#11387 )	2021-04-30 08:08:15 -04:00
Nicolas Patry	db9dd09cf9	Adding `AutomaticSpeechRecognitionPipeline`. (#11337 ) * Adding `AutomaticSpeechRecognitionPipeline`. - Because we added everything to enable this pipeline, we probably should add it to `transformers`. - This PR tries to limit the scope and focuses only on the pipeline part (what should go in, and out). - The tests are very specific for S2T and Wav2vec2 to make sure both architectures are supported by the pipeline. We don't use the mixin for tests right now, because that requires more work in the `pipeline` function (will be done in a follow up PR). - Unsure about the "helper" function `ffmpeg_read`. It makes a lot of sense from a user perspective, it does not add any additional dependencies (as in hard dependency, because users can always use their own load mechanism). Meanwhile, it feels slightly clunky to have so much optional preprocessing. - The pipeline is not done to support streaming audio right now. Future work: - Add `automatic-speech-recognition` as a `task`. And add the FeatureExtractor.from_pretrained within `pipeline` function. - Add small models within tests - Add the Mixin to tests. - Make the logic between ForCTC vs ForConditionalGeneration better. * Update tests/test_pipelines_automatic_speech_recognition.py Co-authored-by: Lysandre Debut <lysandre@huggingface.co> * Adding docs + main import + type checking + LICENSE. * Doc style !. * Fixing TYPE_HINT. * Specifying waveform shape in the docs. * Adding asserts + specify in the documentation the shape of the input np.ndarray. * Update src/transformers/pipelines/automatic_speech_recognition.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Adding require to tests + move the `feature_extractor` doc. Co-authored-by: Lysandre Debut <lysandre@huggingface.co> Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>	2021-04-30 11:54:08 +02:00
CeShine Lee	76116f479b	T5 Gradient Checkpointing (#11353 ) * Implement gradient checkpoinging for T5Stack * A bit more robust type checking * Add `gradient_checkpointing` to T5Config * Formatting * Set requires_grad only when training * None return value will only cause problems when training * Change the output tuple according to `use_cache` * Enable gradient checkpointing for the decoder Squashed commit of the following: commit 658bdd0bd1215353a8770f558bda2ea69a0ad0c7 Author: Ceshine Lee <shuanck@gmail.com> Date: Sat Apr 24 14:08:17 2021 +0800 Only set `require_grad` for gradient checkpointing commit acaeee6b2e675045fb28ce2176444c1d63e908bd Author: Ceshine Lee <shuanck@gmail.com> Date: Sat Apr 24 13:59:35 2021 +0800 Make gradient checkpointing work with the decoder * Formatting	2021-04-30 14:13:55 +05:30
Manuel Romero	58c789e3d2	Update README.md (#11489 ) Add link to code	2021-04-30 04:29:59 -04:00
Patrick von Platen	022a1e9e67	make style (#11520 )	2021-04-30 09:54:58 +02:00
Philip May	e0db8276a6	add sp_model_kwargs to unpickle of xlm roberta tok (#11430 ) add test for pickle simplify test fix test code style add missing pickle import fix test fix test fix test	2021-04-30 03:44:58 -04:00
Frederik Bode	b43e3f93ac	correct the dimension comment of matrix multiplication (#11494 ) Co-authored-by: Frederik Bode <frederik@paperbox.ai>	2021-04-30 09:42:13 +02:00
Lysandre Debut	f37f2adb68	Pin HuggingFace Hub dependency (#11502 )	2021-04-30 02:57:50 -04:00
Lysandre	60d5bda4fd	Patch notification service	2021-04-30 08:56:18 +02:00
Sylvain Gugger	b29eb247d3	Split checkpoint from model_name_or_path in examples (#11492 ) * Split checkpoint from model_name_or_path in examples * Address review comments * Address review comments	2021-04-29 18:33:47 -04:00
Michael Benayoun	d6ec54ba36	solved coefficient issue for the TF version of gelu_fast (#11514 ) Co-authored-by: Michael Benayoun <michael@huggingface.co>	2021-04-29 21:47:26 +02:00
Sylvain Gugger	ad1f7bef13	Reformat to make code clearer in tokenizer call (#11497 ) * Reformat to make code clearer * Reformat to make code clearer	2021-04-29 07:51:09 -04:00
Patrick von Platen	f748bd4242	[Flax] Add docstrings & model outputs (#11498 ) * add attentions & hidden states * add model outputs + docs * finish docs * finish tests * finish impl * del @ * finish * finish * correct test * apply sylvains suggestions * Update src/transformers/models/bert/modeling_flax_bert.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * simplify more Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2021-04-29 12:04:51 +02:00
Hamel Husain	3f6add8bab	fix #1149 (#11493 )	2021-04-28 11:16:41 -04:00
Hamel Husain	c0eb218a55	Update `PreTrainedTokenizerBase` to check/handle batch length for `text_pair` parameter (#11486 ) * Update tokenization_utils_base.py * add assertion * check batch len * Update src/transformers/tokenization_utils_base.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * add error message Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2021-04-28 10:11:17 -04:00
Sylvain Gugger	2d27900b5d	Update min versions in README and add Flax (#11472 ) * Update min versions in README and add Flax * Adapt index	2021-04-28 09:10:06 -04:00
Suraj Patil	8d43c71a1c	fix docs for decoder_input_ids (#11466 ) * fix docs for decoder_input_ids * revert the changes for bart and mbart	2021-04-27 19:36:36 +05:30
Hamel Husain	7ceff67e1a	Finish Making Quick Tour respect the model object (#11467 ) * finish quicktour * fix import * fix print * explain config default better * Update docs/source/quicktour.rst Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2021-04-27 10:04:12 -04:00
Hamel Husain	88ac60f7b5	update QuickTour docs to reflect model output object (#11462 ) * update docs to reflect model output object * run make style`	2021-04-26 22:18:37 -04:00
Ashwin Geet D'Sa	741d48f5c7	Remove max length beam scorer (#11378 ) * removed max_len * removed max_length from BeamSearchScorer * correct max length * finish * del vim * finish & add test Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>	2021-04-27 00:28:40 +02:00
Stas Bekman	bc2571e61c	[Deepspeed] ZeRO-Infinity integration plus config revamp (#11418 ) * adding Z-inf * revamp config process * up version requirement * wip * massive rewrite * cleanup * cleanup * Apply suggestions from code review Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * consistent json commas * act on suggestions * leave this feature for 0.3.16 * style Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2021-04-26 10:40:32 -07:00
Jaimeen Ahn	0661abc545	Variable Correction for Consistency in Distillation Example (#11444 ) As the error comes from the inconsistency of variable meaning number of gpus in parser and its actual usage in the train.py script, 'gpus' and 'n_gpu' respectively, the correction makes the example work	2021-04-26 13:30:48 -04:00
Bhadresh Savani	1d30ec95c7	[Examples] Fixes inconsistency around eval vs val and predict vs test (#11380 ) * added changes for uniformity * modified files * corrected typo * fixed qa scripts * fix typos * fixed predict typo in qa no trainer * fixed test file * reverted trainer changes * reverted trainer changes in custom exmaples * updated readme * added changes in deepspeed test * added changes for predict and eval	2021-04-26 09:24:31 -07:00
Sylvain Gugger	7959d83599	Give each test a different repo name (#11453 )	2021-04-26 11:52:23 -04:00
Sylvain Gugger	b03b2a653d	Style	2021-04-26 11:45:04 -04:00
Stas Bekman	ce11318e7e	make sure to test against the local checkout (#11437 )	2021-04-26 08:42:43 -07:00
Stas Bekman	a753cafdc0	[docs] fix invalid class name (#11438 ) * fix invalid class name * proper ref * proper ref	2021-04-26 08:37:32 -07:00
Kostas Stathoulopoulos	6715e3b6a1	Clarify description of the is_split_into_words argument (#11449 ) * Improve documentation for is_split_into_words argument * Change description wording	2021-04-26 11:29:36 -04:00
Sylvain Gugger	ab2cabb964	Pass along seed to DistributedSampler (#11406 ) * Pass along seed to DistributedSampler * Add seed to DistributedLengthGroupedSampler	2021-04-26 10:26:52 -04:00
LSinev	b24ead87e1	fix some typos in docs, comments, logging/errors (#11432 )	2021-04-26 09:14:25 -04:00
Amine Abdaoui	e3e70f9551	docs(examples): fix link to TPU launcher script (#11427 )	2021-04-26 09:08:43 -04:00

1 2 3 4 5 ...

7165 Commits