transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-30 01:32:23 +06:00

Author	SHA1	Message	Date
Rémi Louf	447fffb21f	process the raw CNN/Daily Mail dataset the data provided by Li Dong et al. were already tokenized, which means that they are not compatible with all the models in the library. We thus process the raw data directly and tokenize them using the models' tokenizers.	2019-10-14 18:12:20 +02:00
Simon Layton	4e6a55751a	Force einsum to fp16	2019-10-14 11:12:41 -04:00
Rémi Louf	67d10960ae	load and prepare CNN/Daily Mail data We write a function to load an preprocess the CNN/Daily Mail dataset as provided by Li Dong et al. The issue is that this dataset has already been tokenized by the authors, so we actually need to find the original, plain-text dataset if we want to apply it to all models.	2019-10-14 14:11:20 +02:00
Timothy Liu	376e65a674	Added automatic mixed precision and XLA options to run_tf_glue.py	2019-10-13 13:19:06 +00:00
Timothy Liu	86f23a1944	Minor enhancements to run_tf_glue.py	2019-10-13 10:21:35 +00:00
VictorSanh	d844db4005	Add citation bibtex	2019-10-11 16:55:42 -04:00
Rémi Louf	b3261e7ace	read parameters from CLI, load model & tokenizer	2019-10-11 18:40:38 +02:00
Rémi Louf	d889e0b71b	add base for seq2seq finetuning	2019-10-11 17:36:12 +02:00
Thomas Wolf	4428aefc63	Merge pull request #1488 from huggingface/pytorch-tpu GLUE on TPU	2019-10-11 16:33:00 +02:00
Luran He	f382a8decd	convert int to str before adding to a str	2019-10-10 19:20:39 -04:00
Lysandre	639f4b7190	Don't save/load when on TPU	2019-10-10 19:17:25 +00:00
Lysandre	d4e7934ac3	GLUE on TPU	2019-10-10 19:03:06 +00:00
Rémi Louf	1e68c28670	add test for initialization of Bert2Rnd	2019-10-10 18:07:11 +02:00
Thomas Wolf	6596e3d566	Merge pull request #1454 from bkkaggle/pytorch-built-in-tensorboard Change tensorboard imports to use built-in tensorboard if available	2019-10-10 11:56:55 +02:00
thomwolf	177a721205	move back to simple space spliting	2019-10-10 11:45:47 +02:00
thomwolf	a5997dd81a	better error messages	2019-10-10 11:31:01 +02:00
Lysandre Debut	2431fea98a	Merge pull request #1383 from keskarnitish/master Adding CTRL	2019-10-09 11:31:05 -04:00
thomwolf	d9e60f4f0d	Merge branch 'master' into pr/1383	2019-10-09 17:25:08 +02:00
Lysandre Debut	e84470ef81	Merge pull request #1384 from huggingface/encoding-qol Quality of life enhancements in encoding + patch MLM masking	2019-10-09 11:18:24 -04:00
jinoobaek-qz	69629c4f0f	Improve naming and only do regex when necessary	2019-10-09 08:48:40 -04:00
jinoobaek-qz	bf34a252b8	Golden path	2019-10-09 08:48:40 -04:00
jinoobaek-qz	528d3f327b	Improve readability and improve make less assumptions about checkpoint format	2019-10-09 08:48:40 -04:00
jinoobaek-qz	56301bd9e8	Extract method	2019-10-09 08:48:40 -04:00
jinoobaek-qz	d6c5469712	Delete older checkpoint after saving new checkpoint	2019-10-09 08:48:40 -04:00
jinoobaek-qz	54a31f50fb	Add save_total_limit	2019-10-09 08:48:40 -04:00
Thomas Wolf	439fac723a	Merge pull request #1409 from brian41005/master Evaluation result.txt path changing #1286	2019-10-09 03:14:34 +02:00
Bilal Khan	5ce8d29abe	Change tensorboard imports to use built-in tensorboard if available	2019-10-08 16:29:43 -05:00
VictorSanh	7ce83b4931	update weights for distilgpt2	2019-10-07 12:30:27 -04:00
LysandreJik	f3e0218fbb	Correct device assignment in run_generation	2019-10-05 21:05:16 -04:00
thomwolf	78ef1a9930	fixes	2019-10-04 17:59:44 -04:00
thomwolf	6c1d0bc066	update encode_plus - add truncation strategies	2019-10-04 17:38:38 -04:00
VictorSanh	0820bb0555	unecessary carriage return	2019-10-04 17:23:15 -04:00
VictorSanh	f5891c3821	run_squad --> run_squad_w_distillation	2019-10-04 17:23:15 -04:00
VictorSanh	764a7923ec	add distillation+finetuning option in run_squad	2019-10-04 17:23:15 -04:00
thomwolf	92c0f2fb90	Merge remote-tracking branch 'origin/julien_multiple-choice' into encoding-qol	2019-10-04 15:48:06 -04:00
Julien Chaumond	9e136ff57c	Honor args.overwrite_cache (h/t @erenup)	2019-10-04 15:00:56 -04:00
keskarnitish	dbed1c5d94	Adding CTRL (squashed commit) adding conversion script adding first draft of modeling & tokenization adding placeholder for test files bunch of changes registering the tokenizer/model/etc tests change link; something is very VERY wrong here weird end-of-word thingy going on i think the tokenization works now ; wrote the unit tests overall structure works;load w next the monster is alive! works after some cleanup as well adding emacs autosave to gitignore currently only supporting the 48 layer one; seems to infer fine on my macbook cleanup fixing some documentation fixing some documentation tests passing? now works on CUDA also adding greedy? adding greedy sampling works well	2019-10-03 22:29:03 -07:00
Lysandre Debut	d3f24dfad7	Merge branch 'master' into master	2019-10-03 22:43:09 +00:00
LysandreJik	ecc4f1bdfa	XLM use_lang_embedding flag in run_generation	2019-10-03 17:42:16 -04:00
LysandreJik	c2c2ca0fdb	Added XLM to run_generation, with prompt language selection.	2019-10-03 17:18:48 -04:00
LysandreJik	aebd83230f	Update naming + remove f string in run_lm_finetuning example	2019-10-03 11:31:36 -04:00
LysandreJik	5ed50a93fb	LM finetuning won't mask special tokens anymore	2019-10-03 11:31:36 -04:00
Brian Ma	7af0777910	Update run_glue.py add DistilBert model shortcut into ALL_MODELS	2019-10-03 15:31:11 +00:00
VictorSanh	5f07d8f11a	prepare release	2019-10-03 10:27:11 -04:00
VictorSanh	35071007cb	incoming release 🔥 update links to arxiv preprint	2019-10-03 10:27:11 -04:00
VictorSanh	2a91f6071f	upddate README - TODO updadte link to paper	2019-10-03 10:27:11 -04:00
VictorSanh	c51e533a5f	update train.py	2019-10-03 10:27:11 -04:00
VictorSanh	a76c3f9cb0	update requirements	2019-10-03 10:27:11 -04:00
VictorSanh	bb9c5ead54	update distiller	2019-10-03 10:27:11 -04:00
VictorSanh	a12ab0a8db	update binarized_data	2019-10-03 10:27:11 -04:00
VictorSanh	4d6dfbd376	update extract	2019-10-03 10:27:11 -04:00
VictorSanh	23edebc079	update extract_distilbert	2019-10-03 10:27:11 -04:00
VictorSanh	cbfcfce205	update token_counts	2019-10-03 10:27:11 -04:00
VictorSanh	19e4ebbe3f	grouped_batch_sampler	2019-10-03 10:27:11 -04:00
VictorSanh	594202a934	lm_seqs_dataset	2019-10-03 10:27:11 -04:00
VictorSanh	38084507c4	add distillation_configs	2019-10-03 10:27:11 -04:00
Brian Ma	2195c0d5f9	Evaluation result.txt path changing #1286	2019-10-03 12:49:12 +08:00
Thomas Wolf	963529e29b	Merge pull request #1288 from echan00/master Typo with LM Fine tuning script	2019-10-01 18:46:07 -04:00
thomwolf	f7978f70ec	use format instead of f-strings	2019-10-01 18:45:38 -04:00
Julien Chaumond	b350662955	overflowing_tokens do not really make sense here, let's just return a number Co-Authored-By: Lysandre Debut <lysandre.debut@reseau.eseo.fr>	2019-09-30 16:37:09 -04:00
Julien Chaumond	f5bcde0b2f	[multiple-choice] Simplify and use tokenizer.encode_plus	2019-09-30 16:04:55 -04:00
Denny	9478590630	Update run_lm_finetuning.py The previous method, just as phrased, did not exist in the class.	2019-09-27 15:18:42 -03:00
Thomas Wolf	d83d295763	Merge pull request #1337 from mgrankin/fastdataset faster dataset building	2019-09-27 10:35:12 +02:00
thomwolf	da2e47ad15	clean up a little run_tf_glue	2019-09-27 09:41:15 +02:00
thomwolf	528c288fa9	clean up run_tf_glue	2019-09-27 09:40:29 +02:00
VictorSanh	702f589848	fix input in run_glue for distilbert	2019-09-27 00:20:14 -04:00
mgrankin	f71a4577b8	faster dataset building	2019-09-26 16:53:13 +03:00
thomwolf	481d9c4fb5	Merge branch 'master' into tf2	2019-09-26 12:02:54 +02:00
thomwolf	31c23bd5ee	[BIG] pytorch-transformers => transformers	2019-09-26 10:15:53 +02:00
thomwolf	5705333441	add initialization for everybody	2019-09-26 10:06:20 +02:00
thomwolf	7c9f8f93f9	fix tests	2019-09-26 01:59:53 +02:00
thomwolf	d6dde438ea	add batch dimension in encode	2019-09-26 01:45:55 +02:00
thomwolf	4a21c4d88d	add warning if neither pt nor tf are found	2019-09-26 01:30:06 +02:00
thomwolf	3b7fb48c3b	fix loading from tf/pt	2019-09-25 17:46:16 +02:00
thomwolf	a049c8043b	push fix to training	2019-09-25 17:33:16 +02:00
mataney	a9f24a16bc	[FIX] fix run_generation.py to work with batch_size > 1	2019-09-25 15:53:29 +03:00
thomwolf	5def3302f4	update run_glue	2019-09-25 12:38:08 +02:00
thomwolf	f71758f7a4	update internal glue processors	2019-09-25 12:00:50 +02:00
thomwolf	b5ec526f85	updated data processor and metrics	2019-09-24 17:10:50 +02:00
LysandreJik	f09e5ecef0	[Proposal] GLUE processors included in library	2019-09-24 09:47:34 -04:00
LysandreJik	c832f43a4d	`output_token_type` -> `token_type_ids`	2019-09-24 07:21:38 -04:00
LysandreJik	3927d7756c	Updated the GLUE pre-processing method	2019-09-24 07:15:11 -04:00
LysandreJik	9d44236f70	Updated DistilBERT	2019-09-24 07:03:24 -04:00
Lorenzo Ampil	4b543c3007	Add option to use a 'stop token' which will be used to truncate the output text to everything till right before the 'stop token'	2019-09-22 21:38:38 +08:00
VictorSanh	9f995b99d4	minor fixes	2019-09-19 21:36:06 +00:00
VictorSanh	3fe5c8e8a8	update bert-base-uncased rslts	2019-09-19 19:34:22 +00:00
VictorSanh	354944e607	[distillation] big update w/ new weights	2019-09-19 19:25:21 +00:00
LysandreJik	60414f31a9	GLUE updated with new methods	2019-09-19 10:55:06 +02:00
LysandreJik	bf503158c5	Sentence -> Sequence. Removed output_mask from the special token addition methods.	2019-09-19 10:55:06 +02:00
LysandreJik	de8e14b6c0	Added DistilBERT to run_squad script	2019-09-19 10:55:06 +02:00
LysandreJik	88368c2a16	Added DistilBERT to `run_lm_finetuning`	2019-09-19 10:55:06 +02:00
LysandreJik	75635072e1	Updated GLUE script to add DistilBERT. Cleaned up unused args in the utils file.	2019-09-19 10:55:06 +02:00
LysandreJik	59057abe52	typo	2019-09-19 10:55:06 +02:00
LysandreJik	bac332fec0	Updated the GLUE data processor. Corrections to RoBERTa and XLNet.	2019-09-19 10:55:06 +02:00
Erik Chan	f0340eccf9	Typo Typo	2019-09-18 13:42:11 -07:00
erenup	8960988f35	fixed to find best dev acc	2019-09-19 01:10:05 +08:00
erenup	46ffc28329	Merge branch 'master' into run_multiple_choice_merge # Please enter a commit message to explain why this merge is necessary, # especially if it merges an updated upstream into a topic branch. # # Lines starting with '#' will be ignored, and an empty message aborts # the commit.	2019-09-18 21:43:46 +08:00
erenup	15143fbad6	move run_multiple_choice.py and utils_multiple_choice.py to examples	2019-09-18 21:18:46 +08:00
erenup	3cd6289758	Merge remote-tracking branch 'huggingface/master' into run_multiple_choice_merge # Conflicts: # examples/contrib/run_swag.py	2019-09-18 21:16:59 +08:00
erenup	36362cf086	move schedule.step after optimizer.step	2019-09-18 21:13:40 +08:00
thomwolf	e768f2322a	update run_openai_gpt to fix #1264	2019-09-18 10:07:47 +02:00
thomwolf	8334993915	clean up examples - updated to new keyword inputs - #1246	2019-09-18 10:01:27 +02:00
erenup	5882c442e5	add example usage	2019-09-16 22:38:08 +08:00
erenup	982f181aa7	Merge remote-tracking branch 'origin/master' into run_multiple_choice_add_doc	2019-09-16 19:12:00 +08:00
erenup	84b9d1c423	Merge remote-tracking branch 'huggingface/master' # Conflicts: # pytorch_transformers/__init__.py	2019-09-16 19:06:12 +08:00
erenup	603b470a3d	add warnning info	2019-09-16 18:53:37 +08:00
erenup	4812a5a767	add doc string	2019-09-16 11:50:18 +08:00
VictorSanh	32e1332acf	[distil] fix once for all general logger for scripts	2019-09-11 14:19:07 +00:00
VictorSanh	364920e216	fix small bug/typo	2019-09-10 21:45:01 +00:00
Thomas Wolf	23c23f5399	Merge pull request #1229 from SKRohit/master changes in evaluate function in run_lm_finetuning.py	2019-09-10 22:16:45 +02:00
searchivarius	eab980fd68	Fix to prevent crashing on assert len(tokens_b)>=1	2019-09-09 19:58:08 -04:00
VictorSanh	a95ced6260	[Distillation] save last chkpt as `pytorch_model.bin`	2019-09-09 19:53:35 +00:00
Rohit Kumar Singh	e5df36397b	changes in return statement of evaluate function changed `results` to `result` and removed `results` dict defined previously	2019-09-09 19:55:57 +05:30
LysandreJik	3f91338be9	Patched a few outdated parameters	2019-09-06 17:48:06 -04:00
LysandreJik	f47f9a5874	Updated outdated examples	2019-09-06 17:10:33 -04:00
LysandreJik	5e151f5e77	Table of contents	2019-09-06 12:08:36 -04:00
LysandreJik	593c070435	Better examples	2019-09-06 12:00:12 -04:00
VictorSanh	dddd6b9927	Update DistilBERT training code	2019-09-05 18:26:14 +00:00
Stefan Schweter	a1c34bd286	distillation: fix ModuleNotFoundError error in token counts script	2019-08-31 12:21:38 +02:00
Thomas Wolf	51e980ce36	Merge pull request #1155 from anhnt170489/apex_fp16 Update apex fp16 implementation	2019-08-30 23:29:11 +02:00
VictorSanh	282c276e09	typos + file name coherence in distillation README	2019-08-30 12:02:29 -04:00
VictorSanh	803c1cc4ea	fix relative import bug cf Issue #1140	2019-08-30 12:01:27 -04:00
Thomas Wolf	0a2fecdf90	Merge branch 'master' into master	2019-08-30 16:30:08 +02:00
Rabeeh KARIMI	39eb31e11e	remove reloading tokenizer in the training, adding it to the evaluation part	2019-08-30 15:44:41 +02:00
Rabeeh KARIMI	350bb6bffa	updated tokenizer loading for addressing reproducibility issues	2019-08-30 15:34:28 +02:00
Thomas Wolf	01ad55f8cf	Merge pull request #1026 from rabeehk/master loads the tokenizer for each checkpoint, to solve the reproducability…	2019-08-30 14:15:36 +02:00
erenup	6e1ac34e2b	Merge remote-tracking branch 'huggingface/master'	2019-08-30 15:50:11 +08:00
jamin	2fb9a934b4	re-format	2019-08-30 14:05:28 +09:00
jamin	c8731b9583	update apex fp16 implementation	2019-08-30 13:54:00 +09:00
LysandreJik	caf1d116a6	Closing bracket in DistilBERT's token count.	2019-08-29 15:30:10 -04:00
Luis	fe8fb10b44	Small modification of comment in the run_glue.py example Add RoBERTa to the comment as it was not explicit that RoBERTa don't use token_type_ids.	2019-08-29 14:43:30 +02:00
erenup	942d3f4b20	modifiy code of arc label insurance	2019-08-29 10:21:17 +08:00
LysandreJik	bf3dc778b8	Changed learning rate for run_squad test	2019-08-28 18:24:43 -04:00
Andreas Daiminger	1d15a7f278	swap order of optimizer.step() and scheduler.step()	2019-08-28 19:18:27 +02:00
Thomas Wolf	0ecfd17f49	Merge pull request #987 from huggingface/generative-finetuning Generative finetuning	2019-08-28 16:51:50 +02:00
thomwolf	b5eb283aaa	update credits	2019-08-28 16:36:55 +02:00
thomwolf	912a377e90	dilbert -> distilbert	2019-08-28 13:59:42 +02:00
thomwolf	4ce5f36f78	update readmes	2019-08-28 12:14:31 +02:00
erenup	ec4b1c659f	logging truth error	2019-08-28 16:50:40 +08:00
erenup	df52abe373	add sep_toekn between question and choice	2019-08-28 16:36:21 +08:00
erenup	43c243254a	avoid invalid labels of truth	2019-08-28 16:03:17 +08:00
erenup	3c7e676f8b	add test related code: test the best dev acc model when model is training	2019-08-28 15:57:29 +08:00
VictorSanh	93e82ab424	Write README for DilBERT	2019-08-28 06:26:09 +00:00
VictorSanh	fea921d382	add licensing	2019-08-28 04:45:39 +00:00
VictorSanh	da1e4e53fc	some fixes in `train.py` for loading previous checkpoint	2019-08-28 04:01:03 +00:00
VictorSanh	0d8f8848d5	add `scripts/extract_for_distil.py`	2019-08-28 04:00:19 +00:00
VictorSanh	7f2c384c80	add `scripts/token_counts.py`	2019-08-28 04:00:03 +00:00
VictorSanh	4d16b279e5	add `scripts/binarized_data.py`	2019-08-28 03:59:48 +00:00
VictorSanh	b247b0d880	add `train.py` for distillation	2019-08-28 02:12:47 +00:00
VictorSanh	780f183e55	add requirements	2019-08-28 01:39:52 +00:00

1 2 3 4 5 ...

622 Commits