transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-31 02:02:21 +06:00

Author	SHA1	Message	Date
Patrick von Platen	7630c11f32	[Wav2Vec2] SpecAugment Fast (#11764 ) * first try * finish	2021-05-25 13:59:52 +01:00
Sylvain Gugger	f086652b16	Add option to log only once in multinode training (#11819 ) * Add option to long only once in multinode training * Use an alternate property	2021-05-25 08:03:43 -04:00
Wang Ran (汪然)	b8344a274f	typo (#11858 )	2021-05-25 04:23:46 -04:00
Shiro T	f9880f62ad	fixed a small typo in the doc (#11856 )	2021-05-25 04:18:55 -04:00
Lysandre Debut	6da129cb31	Enable memory metrics in tests that need it (#11859 )	2021-05-25 04:06:19 -04:00
Lysandre Debut	db0b2477cc	Add some tests to the slow suite #11860	2021-05-25 04:06:06 -04:00
Sylvain Gugger	afe479adb5	[Trainer] Report both steps and num samples per second (#11818 ) * [Trainer] Report both steps and num samples per second * Fix batch number * Update src/transformers/trainer_utils.py Co-authored-by: Stas Bekman <stas00@users.noreply.github.com> * Address review comments Co-authored-by: Stas Bekman <stas00@users.noreply.github.com>	2021-05-24 19:51:42 -04:00
Nick Lane-Smith	eaab9397cd	Fix two typos in docs (#11852 ) * typo2 * fix typo	2021-05-24 14:26:02 -04:00
Teven	8a2a3a25af	Fix flos single node (#11844 ) * fixing flos bug/typo in non-distributed setting * storing flos every logging_interval	2021-05-24 20:15:52 +02:00
Sylvain Gugger	adb785b0fe	Switch mem metrics flag (#11851 ) * Switch mem metrics flag * Update src/transformers/training_args.py Co-authored-by: Stas Bekman <stas00@users.noreply.github.com> Co-authored-by: Stas Bekman <stas00@users.noreply.github.com>	2021-05-24 13:30:39 -04:00
Sylvain Gugger	fcdb85e9d2	Fix reference to XLNet (#11846 )	2021-05-24 09:26:40 -04:00
Patrick von Platen	f580604157	[Flax] Fix PyTorch import error (#11839 ) * fix_torch_device_generate_test * remove @ * change pytorch import to flax import	2021-05-24 10:41:10 +01:00
Lysandre Debut	0cbddfb190	Replace double occurrences as the last step (#11367 )	2021-05-24 03:38:59 -04:00
ctheodoris	73fde1defe	Faster list concat for trainer_pt_utils.get_length_grouped_indices() (#11825 ) get_length_grouped_indices() in LengthGroupedSampler and DistributedLengthGroupedSampler is prohibitively slow for large number of megabatches (in test case takes hours for ~270k megabatches with 100 items each) due to slow list concatenation with sum(megabatches, []). Resolves: #11795 Co-authored-by: ctheodoris <cvtheodo@ds.dfci.harvard.edu>	2021-05-22 10:27:20 -04:00
Patrick von Platen	da22245ed9	Add flax text class colab (#11824 ) * fix_torch_device_generate_test * remove @ * add flax glue link	2021-05-21 23:11:58 +01:00
Stas Bekman	a26f4d6208	[Deepspeed] support `zero.Init` in `from_config` (#11805 ) * support zero.Init in from_config * no need for eval test	2021-05-21 09:07:46 -07:00
Patrick von Platen	82335185fe	[Flax] Small fixes in `run_flax_glue.py` (#11820 ) * fix_torch_device_generate_test * remove @ * correct best seed for flax fine-tuning Co-authored-by: Patrick von Platen <patrick@huggingface.co>	2021-05-21 16:52:23 +01:00
Sylvain Gugger	b8697bc622	Avoid TensorFlow import in Trainer	2021-05-21 09:23:31 -04:00
yujun	e2c1dd0966	fix roformer config doc (#11813 )	2021-05-21 08:06:11 -04:00
Lysandre Debut	1b652295c5	Patch recursive import (#11812 )	2021-05-21 06:50:01 -04:00
Patrick von Platen	bd9871657b	[Flax] Align GLUE training script with mlm training script (#11778 ) * speed up flax glue * remove unnecessary line * remove folder * remove run in loop Co-authored-by: Patrick von Platen <patrick@huggingface.co>	2021-05-21 09:36:56 +01:00
Keren Fuentes	223943872e	Fix failing test on Windows Platform (#11589 ) * add separator for windows * fixes test_is_copy_consistent on Windows * fixing writing encoding issue on extended test (for Windows) * resolving comments	2021-05-20 19:54:23 -04:00
Michael Benayoun	f4a0d6ff86	A cleaner and more scalable implementation of symbolic tracing (#11763 ) Cleaner and more scalable implementation of symbolic tracing with torch.fx, and provides support for new architectures: - ALBERT - DistilBERT - MobileBERT - MegatronBERT - GPT2 - GPT Neo Co-authored-by: Michael Benayoun <michael@huggingface.co>	2021-05-20 18:02:29 +02:00
Sylvain Gugger	469384a777	Fix regression in regression (#11785 ) * Fix regression in regression * Add test	2021-05-20 09:55:13 -04:00
Sylvain Gugger	5ad5cc7198	Fix pattern in conf.py (#11784 )	2021-05-20 09:30:31 -04:00
yujun	206f06f2dd	Add new model RoFormer (use rotary position embedding ) (#11684 ) * add roformer * Update docs/source/model_doc/roformer.rst Co-authored-by: Suraj Patil <surajp815@gmail.com> * Update docs/source/model_doc/roformer.rst Co-authored-by: Suraj Patil <surajp815@gmail.com> * update * add TFRoFormerSinusoidalPositionalEmbedding and fix TFMarianSinusoidalPositionalEmbedding * update docs * make style and make quality * roback * unchanged * rm copies from , this is a error in TFMarianSinusoidalPositionalEmbedding * update Copyright year * move # Add modeling imports here to the correct position * max_position_embeddings can be set to 1536 * # Copied from transformers.models.bert.modeling_bert.BertOutput with Bert->RoFormer * # Copied from transformers.models.bert.modeling_bert.BertLayer.__init__ with Bert->RoFormer * update tokenization_roformer * make style * add staticmethod apply_rotary_position_embeddings * add TF staticmethod apply_rotary_position_embeddings * update torch apply_rotary_position_embeddings * fix tf apply_rotary_position_embeddings error * make style * add pytorch RoFormerSelfAttentionRotaryPositionEmbeddingTest * add TF rotary_position_embeddings test * update test_modeling_rofomer * Update docs/source/model_doc/roformer.rst Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/__init__.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/__init__.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/__init__.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/__init__.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/models/roformer/convert_roformer_original_tf_checkpoint_to_pytorch.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/models/roformer/modeling_roformer.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/models/roformer/modeling_roformer.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/models/roformer/modeling_tf_roformer.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * refact roformer tokenizer * add RoFormerTokenizerFast * add RoFormerTokenizationTest * add require_jieba * update Copyright * update tokenizer & add copy from * add option rotary_value * use rust jieba * use rjieba * use rust jieba * fix test_alignement_methods * slice normalized_string is too slow * add config.embedding_size when embedding_size!=hidden_size * fix pickle tokenizer * Update docs/source/model_doc/roformer.rst Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * make style and make quality Co-authored-by: Suraj Patil <surajp815@gmail.com> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>	2021-05-20 08:00:34 -04:00
Lysandre Debut	075fdab4fe	Deprecate commands from the transformers-cli that are in the hf-cli (#11779 )	2021-05-20 03:16:03 -04:00
Albert Villanova del Moral	2582e59a57	Add DOI badge to README (#11771 )	2021-05-19 09:48:56 -04:00
Patrick von Platen	00440e350f	[Flax MLM] Refactor run mlm with optax (#11745 ) * refactor * update * update * update * refactor run mlm * finalize * refactor more * fix typo * update * finish refactor * modify run mlm * Apply suggestions from code review * Apply suggestions from code review * Apply suggestions from code review * small fixes * upload * upload * finish run mlm script Co-authored-by: Patrick von Platen <patrick@huggingface.co>	2021-05-19 12:00:58 +01:00
Patrick von Platen	43891be19b	[T5 failing CI] Fix generate test (#11770 ) * fix_torch_device_generate_test * remove @	2021-05-19 05:31:17 -04:00
Daniel Stancl	680d181ce8	Fix usage of head masks by PT encoder-decoder models' `generate()` function (#11621 ) * Add missing head masking for generate() function * Add head_mask, decoder_head_mask and cross_attn_head_mask into prepare_inputs_for_generation for generate() function for multiple encoder-decoder models. * Add test_genereate_with_head_masking * [WIP] Update the new test and handle special cases * make style * Omit ProphetNet test so far * make fix-copies	2021-05-19 00:44:53 +01:00
Suraj Patil	ca33278fdb	FlaxGPT2 (#11556 ) * flax gpt2 * combine masks * handle shared embeds * add causal LM sample * style * add tests * style * fix imports, docs, quality * don't use cache * add cache * add cache 1st version * make use cache work * start adding test for generation * finish generation loop compilation * rewrite test * finish * update * update * apply sylvains suggestions * update * refactor * fix typo Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>	2021-05-18 22:50:51 +01:00
Tomy Hsieh	eb3e072a3b	Fix a small error in summarization example (#11762 )	2021-05-18 14:38:36 -04:00
Avital Oliver	77f9bd18af	Add Flax Examples and Cloud TPU README (#11753 ) * Add Flax Examples README * Apply suggestions from code review * Update examples/flax/README.md * add nice table * fix * fix * apply suggestions * upload * finish flax readme.md Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>	2021-05-18 17:45:16 +01:00
Philipp Schmid	04e25c6286	add `dataset_name` to data_args and added accuracy metric (#11760 ) * add `dataset_name` to data_args and added accuracy metric * added documentation for dataset_name * spelling correction	2021-05-18 16:27:29 +02:00
Vyom Pathak	fd3b12e8c3	Fixed: Better names for nlp variables in pipelines' tests and docs. (#11752 ) * Fixed: Better names for nlp variables in pipelines' tests and docs. * Fixed: Better variable names	2021-05-18 09:47:28 -04:00
Patrick von Platen	cebb96f53a	Add more subsections to main doc (#11758 ) * add headers to main doc * Apply suggestions from code review * update * upload	2021-05-18 14:38:56 +01:00
Tommy Chiang	da7e73b721	Fix incorrect newline in #11650 (#11757 )	2021-05-18 15:28:13 +02:00
Sylvain Gugger	a515caa331	Fix checkpoint deletion (#11748 )	2021-05-18 07:42:39 -04:00
Nicolas Patry	b88e0e016d	[TokenClassification] Label realignment for subword aggregation (#11680 ) * [TokenClassification] Label realignment for subword aggregation Tentative to replace https://github.com/huggingface/transformers/pull/11622/files - Added `AggregationStrategy` - `ignore_subwords` and `grouped_entities` arguments are now fused into `aggregation_strategy`. It makes more sense anyway because `ignore_subwords=True` with `grouped_entities=False` did not have a meaning anyway. - Added 2 new ways to aggregate which are MAX, and AVERAGE - AVERAGE requires a bit more information than the others, for now this case is slightly specific, we should keep that in mind for future changes. - Testing has been modified to reflect new argument, and to check the correct deprecation and the new aggregation_strategy. - Put the testing argument and testing results for aggregation_strategy, close together, so that readers can understand what is supposed to happen. - `aggregate` is now only tested on a small model as it does not mean anything to test it globally for all models. - Previous tests are unchanged in desired output. - Added a new test case that showcases better the difference between the FIRST, MAX and AVERAGE strategies. * Wrong framework. * Addressing three issues. 1- Tags might not follow B-, I- convention, so any tag should work now (assumed as B-TAG) 2- Fixed an issue with average that leads to a substantial code change. 3- The testing suite was not checking for the "index" key for "none" strategy. This is now fixed. The issue is that "O" could not be chosen by AVERAGE strategy because those tokens were filtered out beforehand, so their relative scores were not counted in the average. Now filtering on ignore_labels will happen at the very end of the pipeline fixing that issue. It's a bit hard to make sure this stays like that because we do not have a end-to-end test for that behavior * Formatting. * Adding formatting to code + cleaner handling of B-, I- tags. Co-authored-by: Francesco Rubbo <rubbo.francesco@gmail.com> Co-authored-by: elk-cloner <rezakakhki.rk@gmail.com> * Typo. Co-authored-by: Francesco Rubbo <rubbo.francesco@gmail.com> Co-authored-by: elk-cloner <rezakakhki.rk@gmail.com>	2021-05-18 09:53:20 +02:00
Patrick von Platen	c73e35323d	push (#11750 )	2021-05-17 19:54:33 +01:00
Sylvain Gugger	936b57158a	Use new evaluation loop in TrainerQA (#11746 )	2021-05-17 10:10:13 -04:00
Patrick von Platen	73893fc771	[BigBird Pegasus] Make tests faster (#11744 ) * improve tests * remove bogus file * make style Co-authored-by: Patrick von Platen <patrick@huggingface.co>	2021-05-17 06:30:53 -04:00
Michael Benayoun	a0531c8a24	fixed shape issue for T5 tracing (#11742 ) Co-authored-by: Michael Benayoun <michael@huggingface.co>	2021-05-17 06:17:31 -04:00
Julien Chaumond	0fc56df5fb	Add visual + link to Premium Support webpage (#11740 ) * Update README.md * Update index.rst	2021-05-17 05:28:56 -04:00
Julien Chaumond	2f88bd9c4c	Remove tapas model card (#11739 )	2021-05-17 04:42:37 -04:00
Marc van Zee	726e953d44	Improvements to Flax finetuning script (#11727 ) * Add Cloud details to README * Flax script and readme updates * Some simplifications of Flax script	2021-05-17 09:26:33 +01:00
Michael Benayoun	86d5fb0b36	Experimental symbolic tracing feature with torch.fx for BERT, ELECTRA and T5 (#11475 ) Symbolic tracing feature for BERT, ELECTRA and T5 Co-authored-by: Michael Benayoun <michael@huggingface.co> Co-authored-by: Stas Bekman <stas@stason.org> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2021-05-14 20:57:30 +02:00
Marc van Zee	94a2348706	Add Cloud details to README (#11706 ) * Add Cloud details to README * Flax script and readme updates	2021-05-14 14:51:25 +01:00
Patrick von Platen	113eaa7575	correct example script (#11726 )	2021-05-14 12:02:57 +01:00

1 2 3 4 5 ...

7234 Commits