transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-31 02:02:21 +06:00

Author	SHA1	Message	Date
Leandro von Werra	740a1574f1	fix link in performance docs (#17419 )	2022-05-25 20:54:43 +02:00
lewtun	284fc6c0bb	Add link to Hub PR docs in model cards (#17421 )	2022-05-25 20:38:56 +02:00
Cookie_thief	35e2d13f3c	Upd AutoTokenizer.from_pretrained doc examples (#17416 )	2022-05-25 11:35:50 -04:00
Animesh Jain	897a8dd89f	Support compilation via Torchdynamo, AOT Autograd, NVFuser (#17308 ) * Support compilation via Torchdynamo, AOT Autograd, NVFuser * Address comments * Lint * Stas comments - missing quality test * Lintere * Quality test * Doc lint * Reset CUDA peak mem * Add CustomTrainer * require a single gpu Co-authored-by: Stas Bekman <stas@stason.org>	2022-05-25 11:16:09 -04:00
Sylvain Gugger	31484afbed	Add test for new model parallelism features (#17401 )	2022-05-25 10:51:27 -04:00
Sylvain Gugger	56b35ce3eb	Make check_init script more robust and clean inits (#17408 )	2022-05-25 07:23:56 -04:00
Sylvain Gugger	bd908e9bb1	Fix README localizer script (#17407 )	2022-05-25 07:23:40 -04:00
Yih-Dar	4d727bd2df	Fix expected value for OPT test `test_inference_no_head` (#17395 ) * Fix expected value * 5e-5 Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-05-25 11:19:06 +02:00
dependabot[bot]	1ef9a1ed4a	Bump tensorflow in /examples/research_projects/decision_transformer (#17400 ) Bumps [tensorflow](https://github.com/tensorflow/tensorflow) from 2.8.0 to 2.8.1. - [Release notes](https://github.com/tensorflow/tensorflow/releases) - [Changelog](https://github.com/tensorflow/tensorflow/blob/master/RELEASE.md) - [Commits](https://github.com/tensorflow/tensorflow/compare/v2.8.0...v2.8.1) --- updated-dependencies: - dependency-name: tensorflow dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2022-05-24 19:36:55 -04:00
Jason Phang	71e602725b	[WIP] Adding GPT-NeoX-20B (#16659 ) * initial * first try * working 20B * 20B tokenizers * Docs * Import fixes for missing classes * Update docs, fixup * black formatting * isort * flake * dummy objects * documentation * Documentation yml * more docs * tweaks for tests * tokenization auto * fix neox tests * test * test * einsum * address PR feedback * Documentation * Update README.md Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/models/gpt_neox/__init__.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/models/gpt_neox/configuration_gpt_neox.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Apply suggestions from code review Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Remove undefined LaTeX syntax * Update to full url to avoid confusion about if that's supposed to refer to the Hub * fix auto * move tests * documentation fix * more doc fixes * test refactor * fix import * fix import * fix import * fix import * fix import * style fixes * More modeling fixes Co-authored-by: Jason Phang <zp489@gr057.hpc.nyu.edu> Co-authored-by: Stella Biderman <stellabiderman@gmail.com> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2022-05-24 09:31:10 -04:00
NielsRogge	374a2f693f	Clean up CLIP tests (#17380 ) Co-authored-by: Niels Rogge <nielsrogge@Nielss-MacBook-Pro.local>	2022-05-24 14:51:26 +02:00
Nicolas Patry	d980929803	Enabling `imageGPT` auto feature extractor. (#16871 ) * Enablign `imageGPT` auto feature extractor. Co-authored-by: ydshieh <ydshieh@users.noreply.github.com> * Small updates. * Update after rebase to use `input_ids` instead of `pixel_values`. Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-05-24 12:30:46 +02:00
NielsRogge	31ee80d556	Add LayoutLMv3 (#17060 ) * Make forward pass work * More improvements * Remove unused imports * Remove timm dependency * Improve loss calculation of token classifier * Fix most tests * Add docs * Add model integration test * Make all tests pass * Add LayoutLMv3FeatureExtractor * Improve integration test + make fixup * Add example script * Fix style * Add LayoutLMv3Processor * Fix style * Add option to add visual labels * Make more tokenizer tests pass * Fix more tests * Make more tests pass * Fix bug and improve docs * Fix import of processors * Improve docstrings * Fix toctree and improve docs * Fix auto tokenizer * Move tests to model folder * Move tests to model folder * change default behavior add_prefix_space * add prefix space for fast * add_prefix_spcae set to True for Fast * no space before `unique_no_split` token * add test to hightligh special treatment of added tokens * fix `test_batch_encode_dynamic_overflowing` by building a long enough example * fix `test_full_tokenizer` with add_prefix_token * Fix tokenizer integration test * Make the code more readable * Add tests for LayoutLMv3Processor * Fix style * Add model to README and update init * Apply suggestions from code review * Replace asserts by value errors * Add suggestion by @ducviet00 * Add model to doc tests * Simplify script * Improve README * a step ahead to fix * Update pair_input_test * Make all tokenizer tests pass - phew * Make style * Add LayoutLMv3 to CI job * Fix auto mapping * Fix CI job name * Make all processor tests pass * Make tests of LayoutLMv2 and LayoutXLM consistent * Add copied from statements to fast tokenizer * Add copied from statements to slow tokenizer * Remove add_visual_labels attribute * Fix tests * Add link to notebooks * Improve docs of LayoutLMv3Processor * Fix reference to section Co-authored-by: SaulLu <lucilesaul.com@gmail.com> Co-authored-by: Niels Rogge <nielsrogge@Nielss-MacBook-Pro.local>	2022-05-24 09:53:45 +02:00
Sylvain Gugger	13541b4aa2	Add support for `device_map="auto"` to OPT (#17382 )	2022-05-23 15:25:51 -04:00
vfbd	71cced8ae3	OPTForCausalLM lm_head input size should be config.word_embed_proj_dim (#17225 )	2022-05-23 21:20:29 +02:00
Sylvain Gugger	56f50590d5	Use Accelerate in `from_pretrained` for big model inference (#17341 ) * Initial work * More or less finished with first draft * Update src/transformers/modeling_utils.py Co-authored-by: Stas Bekman <stas00@users.noreply.github.com> * Update src/transformers/modeling_utils.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Fix randomly initialized weights * Update src/transformers/modeling_utils.py Co-authored-by: Lysandre Debut <lysandre.debut@reseau.eseo.fr> * Address review comments * Rename DeepSpeed folder to temporarily fix the test issue? * Revert to try if Accelerate fix works * Use latest Accelerate release * Quality and fixes * Style * Quality * Add doc * Test + fix * More blocks Co-authored-by: Stas Bekman <stas00@users.noreply.github.com> Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> Co-authored-by: Lysandre Debut <lysandre.debut@reseau.eseo.fr>	2022-05-23 14:32:21 -04:00
Michael Benayoun	2e7e4280aa	Traced models serialization and torchscripting fix (#17206 ) * Fix torch.jit.script and pickling issues * Fix get_attr issues * Fix import in function * Fix GPT-J and T5 tracing for torch=1.11 * Gate graph surgery on torch version * Modeling minor changes to enable TorchScripting * Model serialization / deserialization test * Remove _assert_is_none users	2022-05-23 17:50:40 +02:00
Maximilian Schmidt	1cd01b0af3	Fix Comet ML integration (#17381 ) Callback function `on_train_end` crashed if Comet ML integration was used but `COMET_MODE` set to `DISABLE`	2022-05-23 10:43:10 -04:00
Anugunj Naman	c86aad6110	Fix cvt docstrings (#17367 )	2022-05-23 16:11:09 +02:00
ghlai9665	7b8cb26953	Correct & Improve Doctests for LayoutLMv2 (#17168 ) * add inference example to LayoutLMv2ForQuestionAnswering, passing doctest * add loss example to LayoutLMv2ForQuestionAnswering, passing doctest * Add correct doctest for LayoutLMv2ForTokenClassification, passing doctest * add correct doctest for LayoutLMv2ForSequenceClassification, passing test * add correct doctest for LayoutLMv2Model, passing test * make fixup * fix to address review comments * make style * fix doctest line break issue, add to documentaiton_tests.txt, address review comments * move comment about layoutlmv2 dependencies to the doc page * format doc page as suggested Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * delete extraneous backtick Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2022-05-23 08:02:31 -04:00
Loubna Ben Allal	b48ac1a094	Fix CodeParrot training script (#17291 ) * average loss over batches and accumulated steps for tracking * fix layernorm weight decay * use AdamW from Pytorch instead of Transformers * add shuffling of sequences inside the batches * add shuffling of sequences inside the batches * add logging dir and reformat code * fix lr tracking * remove Mistral scaling * keep Mistral scaling * reformat code * fix error * fix error * use shuffling function from Pytorch * remove argument for shuffling batch sequences as it isn't optional * update package versions and install accelerate from source * remove unused package * Update loss average over accumulated steps Co-authored-by: Leandro von Werra <lvwerra@users.noreply.github.com> * Update loss average over accumulated steps Co-authored-by: Leandro von Werra <lvwerra@users.noreply.github.com> * use one shuffle buffer argument * compute avg_loss in one line Co-authored-by: Loubna ben allal <loubnabenallal@gmail.com> Co-authored-by: Leandro von Werra <lvwerra@users.noreply.github.com>	2022-05-23 12:55:35 +02:00
Daniel Stancl	b9bb417324	Fix a typo relative_postion_if_large -> relative_position_if_large (#17366 )	2022-05-20 18:41:12 +02:00
Sylvain Gugger	3fd7de49f4	Pin dill to fix examples (#17368 ) * Pin dill for now * Try this version? * force install * Actually use dep in testing * Try a larger pin	2022-05-20 11:00:58 -04:00
Patrick von Platen	54192058f3	[Test OPT] Add batch generation test opt (#17359 ) * up * up	2022-05-19 23:46:26 +02:00
ddobokki	48c22691e3	Fix bug in Wav2Vec2 pretrain example (#17326 )	2022-05-19 22:42:44 +02:00
Nathan Dahlberg	5d6feecf16	fix for 17292 (#17293 )	2022-05-19 22:21:19 +02:00
Patrick von Platen	518bd02c9b	[Generation] Fix Transition probs (#17311 ) * [Draft] fix transition probs * up * up * up * make it work * fix * finish * update	2022-05-19 22:17:02 +02:00
Patrick von Platen	e8714c0307	[OPT] Run test in lower precision on GPU (#17353 ) * [OPT] Run test only in half precision * up * up * up * up * finish * fix on GPU * Update tests/models/opt/test_modeling_opt.py	2022-05-19 22:15:36 +02:00
Nicolas Patry	2b282296f1	Adding `batch_size` test to QA pipeline. (#17330 )	2022-05-19 14:28:12 -04:00
Nicolas Patry	a4386d7e40	[BC] Fixing usage of text pairs (#17324 ) * [BC] Fixing usage of text pairs The BC is actually preventing users from misusing the pipeline since users could have been willing to send text pairs and the pipeline would instead understand the thing as a batch returning bogus results. The correct usage of text pairs is preserved in this PR even when that makes the code clunky. Adds support for {"text":..,, "text_pair": ...} inputs for both dataset iteration and more explicit usage to pairs. * Updating the doc. * Update src/transformers/pipelines/text_classification.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update src/transformers/pipelines/text_classification.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Update tests/pipelines/test_pipelines_text_classification.py Co-authored-by: Lysandre Debut <lysandre@huggingface.co> * quality. Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: Lysandre Debut <lysandre@huggingface.co>	2022-05-19 10:29:16 +02:00
Stas Bekman	3601aa8fc9	[tests] fix copy-n-paste error (#17312 ) * [tests] fix copy-n-paste error * fix	2022-05-18 16:00:47 -07:00
Yih-Dar	1b20c970a2	Fix ci_url might be None (#17332 ) * fix * Update utils/notification_service.py Co-authored-by: Lysandre Debut <lysandre.debut@reseau.eseo.fr> Co-authored-by: ydshieh <ydshieh@users.noreply.github.com> Co-authored-by: Lysandre Debut <lysandre.debut@reseau.eseo.fr>	2022-05-18 21:49:08 +02:00
Yih-Dar	6aad3872ce	fix (#17337 ) Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-05-18 15:26:44 -04:00
Zachary Mueller	1762ded30a	Fix metric calculation in examples and setup tests to run on multi-gpu for no_trainer scripts (#17331 ) * Fix length in no_trainer examples * Add setup and teardown * Use new accelerator config generator to automatically make tests able to run based on environment	2022-05-18 14:17:40 -04:00
Jader Martins	6e195eb9de	docs for typical decoding (#17186 ) Co-authored-by: Jader Martins <jadermcs94@gmail.com>	2022-05-18 19:18:43 +02:00
Yih-Dar	060fe61dff	Not send successful report (#17329 ) * send report only if there is any failure Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-05-18 19:07:48 +02:00
Yih-Dar	b3b9f99ed2	Fix test_t5_decoder_model_past_large_inputs (#17320 ) Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-05-18 17:57:23 +02:00
Jingya HUANG	6da76b9c2a	Add onnx export cuda support (#17183 ) Co-authored-by: Lysandre Debut <lysandre@huggingface.co> Co-authored-by: lewtun <lewis.c.tunstall@gmail.com>	2022-05-18 17:52:13 +02:00
NielsRogge	adc0ff2502	Add CvT (#17299 ) * Adding cvt files * Adding cvt files * changes in init file * Adding cvt files * changes in init file * Style fixes * Address comments from code review * Apply suggestions from code review Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> * Format lists in docstring * Fix copies * Apply suggestion from code review Co-authored-by: AnugunjNaman <anugunjjha@gmail.com> Co-authored-by: Ayushman Singh <singhayushman13@protonmail.com> Co-authored-by: Niels Rogge <nielsrogge@Nielss-MacBook-Pro.local> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2022-05-18 17:47:18 +02:00
Sylvain Gugger	4710702837	Fix style	2022-05-18 10:46:40 -04:00
mraunak	5fdb54ece7	Add Information Gain Filtration algorithm (#16953 ) * Add information gain filtration algorithm * Complying with black requirements * Added author * Fixed import order * flake8 corrections Co-authored-by: Javier Turek <javier.turek@intel.com>	2022-05-18 10:39:02 -04:00
Kamal Raj	91ede485a7	Fix typo (#17328 )	2022-05-18 10:29:53 -04:00
Yih-Dar	fe28eb9452	remove (#17325 ) Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-05-18 10:06:41 -04:00
Nicolas Patry	2cb2ea3fa1	Accepting real pytorch device as arguments. (#17318 ) * Accepting real pytorch device as arguments. * is_torch_available.	2022-05-18 10:06:24 -04:00
Nicolas Patry	1c9d1f4ca8	Updating the docs for `max_seq_len` in QA pipeline (#17316 )	2022-05-18 15:46:12 +02:00
Patrick von Platen	60ad73448c	[T5] Fix init in TF and Flax for pretraining (#17294 ) * fix init * Apply suggestions from code review * fix * finish * Update src/transformers/modeling_tf_utils.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2022-05-18 15:08:56 +02:00
Joaq	7ba1d4e51f	Add type hints for ProphetNet (Pytorch) (#17223 ) * added type hints to prophetnet * reformatted with black * fix bc black misformatted some parts * fix imports * fix imports * Update src/transformers/models/prophetnet/configuration_prophetnet.py Co-authored-by: Matt <Rocketknight1@users.noreply.github.com> * update OPTIONAL type hint and docstring Co-authored-by: Matt <Rocketknight1@users.noreply.github.com>	2022-05-18 13:23:47 +01:00
Carl	d6b8e9cec7	Add trajectory transformer (#17141 ) * Add trajectory transformer Fix model init Fix end of lines for .mdx files Add trajectory transformer model to toctree Add forward input docs Fix docs, remove prints, simplify prediction test Apply suggestions from code review Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Apply suggestions from code review Co-authored-by: Lysandre Debut <lysandre@huggingface.co> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Update docs, more descriptive comments Apply suggestions from code review Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Update readme Small comment update and add conversion script Rebase and reformat Fix copies Fix rebase, remove duplicates Fix rebase, remove duplicates * Remove tapex * Remove tapex * Remove tapex	2022-05-17 19:07:43 -04:00
Patrick von Platen	c35264007b	fix (#17310 )	2022-05-17 18:34:31 -04:00
Cesare Campagnano	d9050dc768	[LED] fix global_attention_mask not being passed for generation and docs clarification about grad checkpointing (#17112 ) * [LED] fixed global_attention_mask not passed for generation + docs clarification for gradient checkpointing * LED docs clarification Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * [LED] gradient_checkpointing=True should be passed to TrainingArguments Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * [LED] docs: remove wrong word Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * [LED] docs fix typo Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>	2022-05-17 23:44:37 +02:00

1 2 3 4 5 ...

9868 Commits