transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-08-02 19:21:31 +06:00

Author	SHA1	Message	Date
Stas Bekman	f6c0680d36	add pl_glue example test (#6034 ) * add pl_glue example test * for now just test that it runs, next validate results of eval or predict? * complete the run_pl_glue test to validate the actual outcome * worked on my machine, CI gets less accuracy - trying higher epochs * match run_pl.sh hparms * more epochs? * trying higher lr * for now just test that the script runs to a completion * correct the comment * if cuda is available, add --fp16 --gpus=1 to cover more bases * style	2020-08-11 03:16:52 -04:00
Pradhy729	b25cec13c5	Feed forward chunking (#6024 ) * Chunked feed forward for Bert This is an initial implementation to test applying feed forward chunking for BERT. Will need additional modifications based on output and benchmark results. * Black and cleanup * Feed forward chunking in BertLayer class. * Isort * add chunking for all models * fix docs * Fix typo Co-authored-by: patrickvonplaten <patrick.v.platen@gmail.com>	2020-08-11 03:12:45 -04:00
Lysandre	8a3db6b303	Add TPU testing once again	2020-08-11 08:49:37 +02:00
zcain117	f65ac1faf2	Add missing docker arg for TPU CI. (#6393 )	2020-08-11 02:48:49 -04:00
Sam Shleifer	b9ecd92ee4	[s2s] Script to save wmt data to disk (#6403 )	2020-08-10 22:49:39 -04:00
Patrick von Platen	00bb0b25ed	TF Longformer (#5764 ) * improve names and tests longformer * more and better tests for longformer * add first tf test * finalize tf basic op functions * fix merge * tf shape test passes * narrow down discrepancies * make longformer local attn tf work * correct tf longformer * add first global attn function * add more global longformer func * advance tf longformer * finish global attn * upload big model * finish all tests * correct false any statement * fix common tests * make all tests pass except keras save load * fix some tests * fix torch test import * finish tests * fix test * fix torch tf tests * add docs * finish docs * Update src/transformers/modeling_longformer.py Co-authored-by: Lysandre Debut <lysandre@huggingface.co> * Update src/transformers/modeling_tf_longformer.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * apply Lysandres suggestions * reverse to assert statement because function will fail otherwise * applying sylvains recommendations * Update src/transformers/modeling_longformer.py Co-authored-by: Sam Shleifer <sshleifer@gmail.com> * Update src/transformers/modeling_tf_longformer.py Co-authored-by: Lysandre Debut <lysandre@huggingface.co> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: Sam Shleifer <sshleifer@gmail.com>	2020-08-10 23:25:06 +02:00
Patrick von Platen	3425936643	[EncoderDecoderModel] add a `add_cross_attention` boolean to config (#6377 ) * correct encoder decoder model * Apply suggestions from code review * apply sylvains suggestions	2020-08-10 19:46:48 +02:00
Sylvain Gugger	06bc347c97	Fix links for open in colab (#6391 )	2020-08-10 11:16:17 -04:00
Sylvain Gugger	3e0fe3cf5c	Colab button (#6389 ) * Add colab button * Add colab link for tutorials	2020-08-10 11:12:29 -04:00
Lysandre Debut	79588e6fdb	Ci GitHub caching (#6382 ) * Cache Github Actions CI * Remove useless file	2020-08-10 10:39:31 -04:00
Lysandre Debut	b99098abc7	Patch models (#6326 ) * TFAlbertFor{TokenClassification, MultipleChoice} * Patch models * BERT and TF BERT info s * Update check_repo	2020-08-10 10:39:17 -04:00
Sylvain Gugger	6028ed92bd	Small docfile fixes (#6328 )	2020-08-10 05:37:12 -04:00
Stas Bekman	1429b920d4	refactor almost identical tests (#6339 ) * refactor almost identical tests * important to add a clear assert error message * make the assert error even more descriptive than the original bt	2020-08-10 05:31:20 -04:00
Rohit Gupta	35eb96de4d	correct pl link in readme (#6364 )	2020-08-10 03:08:46 -04:00
Stas Bekman	0830e79512	the test now works again (#6371 )	2020-08-10 02:55:52 -04:00
Alexander Measure	3a556b0fb7	Update modeling_tf_utils.py (#6372 ) fix typo: ckeckpoint->checkpoint	2020-08-10 02:55:11 -04:00
Lysandre	1bbc54a87c	Temporarily de-activate TPU CI	2020-08-10 08:11:40 +02:00
M. Yusuf Sarıgöz	6e8a38568e	[model_cards] electra-base-turkish-cased-ner (#6350 ) * for electra-base-turkish-cased-ner * Add metadata Co-authored-by: Julien Chaumond <chaumond@gmail.com>	2020-08-09 03:39:51 -04:00
Sam Shleifer	9a5ef83748	[s2s] fix --gpus clarg collision (#6358 )	2020-08-08 21:51:37 -04:00
Patrick von Platen	1aec991643	[GPT2] Correct typo in docs (#6352 )	2020-08-08 20:37:29 +02:00
elsanns	9f57e39f71	Add notebook on fine-tuning and interpreting Electra (#6321 ) Co-authored-by: eliska <3648991+elisans@users.noreply.github.com>	2020-08-08 11:47:33 +02:00
Suraj Patil	9bed355449	[s2s] fix label_smoothed_nll_loss (#6344 )	2020-08-08 04:21:12 -04:00
Sam Shleifer	99f73bcc71	[s2s] tiny QOL improvement: run_eval prints scores (#6341 )	2020-08-08 02:45:55 -04:00
Stas Bekman	322dffc6c9	remove a TODO item to use a tiny model (#6338 ) as discussed with @sshleifer, removing this TODO to switch to a tiny model, since it won't be able to test the results of the evaluation (i.e. the results are meaningless).	2020-08-07 21:30:39 -04:00
Sam Shleifer	1f8e826518	[CI] Self-scheduled runner also pins torch (#6332 )	2020-08-07 18:40:21 -04:00
zcain117	1b8a7ffcfd	Add setup for TPU CI to run every hour. (#6219 ) * Add setup for TPU CI to run every hour. * Re-organize config.yml Co-authored-by: Lysandre <lysandre.debut@reseau.eseo.fr>	2020-08-07 11:17:07 -04:00
Stas Bekman	6695450a23	[examples] consistently use --gpus, instead of --n_gpu (#6315 )	2020-08-07 10:36:32 -04:00
Julien Plu	0e36e51515	Fix the tests for Electra (#6284 ) * Fix the tests for Electra * Apply style	2020-08-07 09:30:57 -04:00
Sylvain Gugger	6ba540b747	Add a script to check all models are tested and documented (#6298 ) * Add a script to check all models are tested and documented * Apply suggestions from code review Co-authored-by: Kevin Canwen Xu <canwenxu@126.com> * Address comments Co-authored-by: Kevin Canwen Xu <canwenxu@126.com>	2020-08-07 09:18:37 -04:00
Stas Bekman	e1638dce16	fix the slow tests doc (#6167 ) remove unnecessary duplication wrt `RUN_SLOW=yes`	2020-08-07 09:17:32 -04:00
Binny Mathew	7e9861f7f4	dehate-bert Model Card (#6248 ) Added citation and paper links.	2020-08-07 17:51:03 +08:00
Binny Mathew	f6df6d98dd	dehate-bert Model Card (#6249 ) Added citation and paper links.	2020-08-07 17:48:38 +08:00
Binny Mathew	26691ecba6	dehate-bert Model Card (#6250 ) Added citation and paper links.	2020-08-07 17:48:09 +08:00
Binny Mathew	60657b295c	dehate-bert Model Card (#6251 ) Added citation and paper links.	2020-08-07 17:47:42 +08:00
Binny Mathew	7218261991	dehate-bert Model Card (#6252 ) Added citation and paper links.	2020-08-07 17:47:26 +08:00
Binny Mathew	396d227cd4	dehate-bert Model Card (#6253 ) Added citation and paper links.	2020-08-07 17:47:04 +08:00
Binny Mathew	8be260f18a	dehate-bert Model Card (#6254 ) Added citation and paper links.	2020-08-07 17:46:27 +08:00
Binny Mathew	dce7278cdf	dehate-bert Model Card (#6255 ) Added citation and paper links.	2020-08-07 17:45:52 +08:00
idoh	3be2d04884	fix consistency CrossEntropyLoss in modeling_bart (#6265 )	2020-08-07 17:44:28 +08:00
Lysandre	c72f9c90a1	Remove --no-cache-dir from github CI	2020-08-07 09:07:22 +02:00
Lysandre Debut	0d9328f2ef	Patch GPU failures (#6281 ) * Pin to 1.5.0 * Patch XLM GPU test	2020-08-07 02:58:15 -04:00
Lysandre Debut	80a0676a51	CI dependency wheel caching (#6287 ) * Single workflow cache test Remove cache dir, re-trigger cache Only pip archives Not sudo when pip * All workflow cache Remove no-cache-dir instruction Remove last sudo occurrences v0.3	2020-08-07 02:48:59 -04:00
Stas Bekman	175cd45e13	fix the shuffle agrument usage and the default (#6307 )	2020-08-06 20:32:28 -04:00
Bhashithe Abeysinghe	ffceef2042	[Fix] text-classification PL example (#6027 ) Co-authored-by: Sam Shleifer <sshleifer@gmail.com>	2020-08-06 15:46:43 -04:00
xujiaze13	eb2bd8d6eb	Remove redundant line in run_pl_glue.py (#6305 )	2020-08-06 15:43:45 -04:00
Patrick von Platen	118ecfd427	fix for pytorch < 1.6 (#6300 )	2020-08-06 21:14:46 +02:00
Sam Shleifer	2804fff839	[s2s]Use prepare_translation_batch for Marian finetuning (#6293 ) Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2020-08-06 14:58:38 -04:00
Teven	2f2aa0c89c	added `n_inner` argument to gpt2 config (#6296 )	2020-08-06 17:47:32 +02:00
Manuel Romero	0a0d53dcf8	Update model card (#6290 ) Add links to RuPERTa models fine-tuned on Spanish SQUAD datasets	2020-08-06 11:42:43 -04:00
Doug Blank	b923871bb7	Adds comet_ml to the list of auto-experiment loggers (#6176 ) * Support for Comet.ml * Need to import comet first * Log this model, not the one in the backprop step * Log args as hyperparameters; use framework to allow fine control * Log hyperparameters with context * Apply black formatting * isort fix integrations * isort fix __init__ * Update src/transformers/trainer.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/trainer.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/trainer_tf.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Address review comments * Style + Quality, remove Tensorboard import test Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: Lysandre <lysandre.debut@reseau.eseo.fr>	2020-08-06 11:31:30 -04:00

1 2 3 4 5 ...

4795 Commits