transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-31 02:02:21 +06:00

Author	SHA1	Message	Date
Nicolas Patry	ec6878f6ca	Now supporting pathlike in pipelines too. (#20030 )	2022-11-03 09:14:45 +01:00
Steven Liu	aa39967b28	reorganize glossary (#20010 )	2022-11-02 16:58:17 -07:00
Yih-Dar	305e8718b4	Show installed libraries and their versions in CI jobs (#20026 ) * Show versions * check * store outputs * revert Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-11-02 20:52:39 +01:00
Ben Eyal	9f9ddcc2de	🚨 🚨 🚨 Fix Issue 15003: SentencePiece Tokenizers Not Adding Special Tokens in `convert_tokens_to_string` (#15775 ) * Add test for SentencePiece not adding special tokens to strings * Add SentencePieceStringConversionMixin to fix issue 15003 * Fix conversion from tokens to string for most SentencePiece tokenizers Tokenizers fixed: - AlbertTokenizer - BarthezTokenizer - CamembertTokenizer - FNetTokenizer - M2M100Tokenizer - MBart50Tokenizer - PegasusTokenizer - Speech2TextTokenizer * Fix MarianTokenizer, adjust SentencePiece test to accomodate vocab * Fix DebertaV2Tokenizer * Ignore LayoutXLMTokenizer in SentencePiece string conversion test * Run 'make style' and 'make quality' * Clean convert_tokens_to_string test Instead of explicitly ignoring LayoutXLMTokenizer in the test, override the test in LayoutLMTokenizationTest and do nothing in it. * Remove commented out code * Improve robustness of convert_tokens_to_string test Instead of comparing lengths of re-tokenized text and input_ids, check that converting all special tokens to string yields a string with all special tokens. * Inline and remove SentencePieceStringConversionMixin The convert_tokens_to_string method is now implemented in each relevant SentencePiece tokenizer. * Run 'make style' and 'make quality' * Revert removal of space in convert_tokens_to_string * Remove redundant import * Revert test text to original * Uncomment the lowercasing of the reverse_text variable * Mimic Rust tokenizer behavior for tokenizers - Albert - Barthez - Camembert - MBart50 - T5 * Fix accidentally skipping test in wrong tokenizer * Add test for equivalent Rust and slow tokenizer behavior * Override _decode in BigBirdTokenizer to mimic Rust behavior * Override _decode in FNetTokenizer to mimic Rust behavior * Override _decode in XLNetTokenizer to mimic Rust behavior * Remove unused 're' import * Update DebertaV2Tokenizer to mimic Rust tokenizer * Deberta tokenizer now behaves like Albert and its `convert_tokens_to_string` is not tested. * Ignore problematic tests in Deberta V2 * Add comment on why the Deberta V2 tests are skipped	2022-11-02 15:45:38 -04:00
Yih-Dar	fb7cbe236b	Fix doctest (#20023 ) * Fix doctest Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-11-02 19:37:25 +01:00
Yih-Dar	f69eb24b5a	Improve model tester (#19984 ) * part 1 * part 2 * part 3 * fix * For CANINE * For ESMFold Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-11-02 17:38:44 +01:00
Saad Mahmud	7487743793	[Doctest] Add configuration_deberta_v2.py (#19995 ) * Add example docstring for DebertaV2Config * Add DebertaV2Config to documentation_tests * Fix mistake with directory name	2022-11-02 16:22:11 +01:00
amyeroberts	9aedce99b0	Update auto processor to check image processor created (#20021 )	2022-11-02 15:19:33 +00:00
Sylvain Gugger	49b77b89ea	Quality (#20002 )	2022-11-02 09:53:37 -04:00
Yih-Dar	c6c9db3d0c	Fix gradient checkpoint test in encoder-decoder (#20017 ) Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-11-02 14:15:09 +01:00
amyeroberts	a6b7759880	Add Image Processors (#19796 ) * Add CLIP image processor * Crop size as dict too * Update warning * Actually use logger this time * Normalize doesn't change dtype of input * Add perceiver image processor * Tidy up * Add DPT image processor * Add Vilt image processor * Tidy up * Add poolformer image processor * Tidy up * Add LayoutLM v2 and v3 imsge processors * Tidy up * Add Flava image processor * Tidy up * Add deit image processor * Tidy up * Add ConvNext image processor * Tidy up * Add levit image processor * Add segformer image processor * Add in post processing * Fix up * Add ImageGPT image processor * Fixup * Add mobilevit image processor * Tidy up * Add postprocessing * Fixup * Add VideoMAE image processor * Tidy up * Add ImageGPT image processor * Fixup * Add ViT image processor * Tidy up * Add beit image processor * Add mobilevit image processor * Tidy up * Add postprocessing * Fixup * Fix up * Fix flava and remove tree module * Fix image classification pipeline failing tests * Update feature extractor in trainer scripts * Update pad_if_smaller to accept tuple and int size * Update for image segmentation pipeline * Update src/transformers/models/perceiver/image_processing_perceiver.py Co-authored-by: Alara Dirik <8944735+alaradirik@users.noreply.github.com> * Update src/transformers/image_processing_utils.py Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com> * Update src/transformers/models/beit/image_processing_beit.py Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com> * PR comments - docstrings; remove accidentally added resize; var names * Update docstrings * Add exception if size is not in the right format * Fix exception check * Fix up * Use shortest_edge in tuple in script Co-authored-by: Alara Dirik <8944735+alaradirik@users.noreply.github.com> Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>	2022-11-02 11:57:36 +00:00
Ripose	2e3452af0f	make sentencepiece import conditional in bertjapanesetokenizer (#20012 )	2022-11-02 07:44:37 -04:00
Yih-Dar	8827e1b217	clean up vision/text config dict arguments (#19954 ) * clean up * For backward compatibility * clean up * Same changes for more models Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-11-02 12:03:43 +01:00
Alara Dirik	cb630ffab8	Update object detection pipeline to use post_process_object_detection methods(#20004 )	2022-11-02 10:26:36 +03:00
Steven Liu	79c720c062	fix typo (#20006 )	2022-11-01 11:30:36 -07:00
Joao Gante	831590f6a9	Generate: contrastive search with full optional outputs (#19963 ) * Use beam search functionality; Add extra outputs and test * Add full tests for contrastive search * Add error message on unconventional cache format	2022-11-01 18:15:36 +00:00
Steven Liu	ab74ac11e4	Add LayoutLMv3 resource (#19932 ) * add layoutlmv3 resource * add layoutlmv2 resources * fix button	2022-11-01 11:10:46 -07:00
Steven Liu	dec8578e70	Add BERT resources (#19852 ) * add resources for bert * add course chapters * apply reviews * add pipeline icons and community resource * fix buttons	2022-11-01 11:09:53 -07:00
Steven Liu	1f6885bad0	add dataset (#20005 )	2022-11-01 10:37:20 -07:00
Matt	4f1e5e4efd	Add ESMFold code sample (#20000 ) * Add ESMFold code sample * sorry sylvain * make fixup * sorry sylvain again	2022-11-01 13:21:12 +00:00
Ikko Ashimine	38e5b71abb	Add Japanese translated README (#19945 ) * Add japanese translated README.md * Add README_ja.md link * Add japanese transkate to check_copies.py * Add guide to Japanese README.md * Update README_ja.md Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com> * Update utils/check_copies.py Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2022-11-01 09:18:08 -04:00
Wang Ran (汪然)	4f90fc1db8	typo (#20001 )	2022-11-01 09:04:53 -04:00
Sayak Paul	c87ae86a8f	Update image_classification.mdx (#19996 )	2022-11-01 07:54:41 -04:00
Mohit Sharma	c796b6dea6	Added onnx config whisper (#19525 ) * Added onnx config whisper * added whisper support onnx * add audio input data * added whisper support onnx * fixed the seqlength value * Updated the whisper onnx ocnfig * restore files to old version * removed attention mask from inputs * Updated get_dummy_input_onnxruntime docstring * Updated relative imports and token generation * update docstring	2022-11-01 07:50:42 -04:00
Sylvain Gugger	c3a93d8d82	v4.25.0.dev0	2022-10-31 21:48:40 -04:00
Matt	7f9b7b3f0e	Add ESMFold (#19977 ) * initial commit * First draft that gets outputs without crashing! * Add all the ported openfold dependencies * testing * Restructure config files for ESMFold * Debugging to find output discrepancies * Mainly style * Make model runnable without extra deps * Remove utils and merge them to the modeling file * Use correct gelu and remove some debug prints * More cleanup * Update esm docs * Update conversion script to support ESMFold properly * Port some top-level changes from ESMFold repo * Expand EsmFold docstrings * Make attention_mask optional (default to all 1s) * Add inference test for ESMFold * Use config and not n kwargs * Add modeling output class * Remove einops * Remove chunking in ESM FFN * Update tests for ESMFold * Quality * REpo consistency * Remove tree dependency from ESMFold * make fixup * Add an error in case my structure map function breaks later * Remove needless code * Stop auto-casting the LM to float16 so CPU tests pass * Stop auto-casting the LM to float16 so CPU tests pass * Final test updates * Split test file * Copyright and quality * Unpin PyTorch to see built doc * Fix config file to_dict() method * Add some docstrings to the output * Skip TF checkpoint tests for ESM until we reupload those * make fixup * More docstrings * Unpin to get even with main * Flag example to write Co-authored-by: Sylvain Gugger <Sylvain.gugger@gmail.com>	2022-10-31 21:32:58 -04:00
NielsRogge	4c9e0f029e	Add support for gradient checkpointing (#19990 ) Co-authored-by: Niels Rogge <nielsrogge@Nielss-MacBook-Pro.local>	2022-10-31 18:37:17 +01:00
Yih-Dar	8214a9f66a	Pin torch to < 1.13 temporarily (#19989 ) * pin torch to < 1.13 * pin torch to < 1.13 Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-10-31 18:22:52 +01:00
Jean Charles Kouame	6aede2d602	Tranformers documentation translation to Italian #17459 (#19988 )	2022-10-31 13:19:15 -04:00
Sanchit Gandhi	f38a145418	[ASR] Update 'tasks' for model card (#19986 )	2022-10-31 16:50:17 +00:00
Sanchit Gandhi	9406c7bc82	[modelcard] Update for ASR (#19985 ) * [modelcard] Update for ASR * style	2022-10-31 16:49:58 +00:00
Chiao	225c36fbe5	gradient checkpointing for GPT-NeoX (#19946 ) * gradient checkpointing for GPT-NeoX * initialize gradient checkpointing flag * must set flag before init	2022-10-31 12:32:46 -04:00
Saad Mahmud	6176e13612	[Doctest] Add configuration_deberta.py (#19968 ) * Add Example docstring to DebertaConfig * Add configuration_deberta to documentation_tests * Add microsoft/deberta-base to example docstring * Fix example docstring mistake	2022-10-31 17:22:01 +01:00
Yih-Dar	b047472650	donut -> donut-swin (#19920 ) * donut -> donut-swin * remove ("donut-swin", "DonutProcessor") Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-10-31 14:56:16 +01:00
Sylvain Gugger	a83bb45fb8	Fix repo consistency	2022-10-31 06:42:46 -04:00
lewtun	243439a827	Fix ONNX tests for ONNX Runtime v1.13.1 (#19950 ) * Fix ONNX tests for ONNX Runtime v1.13.1 Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2022-10-31 09:21:45 +01:00
NielsRogge	0b294c2334	[Conditional, Deformable DETR] Add postprocessing methods (#19709 ) * Add postprocessing methods * Update docs * Add fix * Add test * Add test for deformable detr postprocessing * Add post processing methods for segmentation * Update code examples * Add post_process to make the pipeline work * Apply updates Co-authored-by: Niels Rogge <nielsrogge@Nielss-MacBook-Pro.local>	2022-10-31 08:28:44 +01:00
Steven Liu	2e35bac4e7	Add wav2vec2 resources (#19931 ) * add wav2vec2 resources * apply review Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com> Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com>	2022-10-28 13:28:18 -07:00
Steven Liu	9d2788b46b	add resources for distilbert (#19930 )	2022-10-28 13:16:07 -07:00
Steven Liu	b0a2c3a2d6	add resources for bart (#19928 )	2022-10-28 13:15:43 -07:00
എതിരാളിക്കൊരു പോരാളി	98c9c5add9	Update Code of Conduct to Contributor Covenant v2.1 (#19935 ) * Update Code of Conduct to Contributor Covenant v2.1 * Update CODE_OF_CONDUCT.md	2022-10-28 11:03:38 -04:00
Raghav Prabhakar	0d4c45c585	Add Onnx Config for ImageGPT (#19868 ) * add Onnx Config for ImageGPT * add generate_dummy_inputs for onnx config * add TYPE_CHECKING clause * Update doc for generate_dummy_inputs Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com> Co-authored-by: Sylvain Gugger <35901082+sgugger@users.noreply.github.com>	2022-10-28 09:39:53 -04:00
Yang Yu	9b1dcba94a	Use self._trial to generate trial_name for Trainer. (#19874 ) * Do not generate trial_name when trail is None * Use (trial or self._trial) to generate trial_name * Follow comments	2022-10-28 08:47:47 -04:00
donguk.lim	347ba38cb4	Support segformer fx (#19924 ) * Support segformer fx * Add fx_compatible attribute to test_modeling_segformer.py * Update glpn model (fx support) glpn model was copied from segformer. * Update utils/fx.py \| add semantic-segmentation for SegformerForSemanticSegmentation model * Fix minor import order(isort) * Add random input generation for segformer fx Co-authored-by: noelbird <lduldu00228@gmail.com>	2022-10-28 08:44:38 -04:00
Yih-Dar	dcca71be61	Create dummy models (#19901 ) * create dummy models * quality * update * update * Make Wav2Vec2Conformer work * style * deal with models with text_config and vision_config * apply suggestions * Composite models * style * style * fix shape issue * fix shape issue * For VisionTextDualEncoderModel * show_progress=False when converting tokenizers * Fix for OwlViT * Fix for VisualBert * Update * final Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2022-10-28 13:05:41 +02:00
Younes Belkada	4cef546ffc	Add `accelerate` support for BART-like models (#19927 ) * forward contrib credits from suggestion * add `accelerate` support for BART-like models Co-authored-by: sgugger <sgugger@users.noreply.github.com>	2022-10-27 23:14:53 +02:00
Sylvain Gugger	ebfd7229d2	Let inputs of fast tokenizers be tuples as well as lists (#19898 ) * Let inputs of fast tokenizers be tuples as well as lists * Update src/transformers/tokenization_utils_fast.py Co-authored-by: Lysandre Debut <lysandre.debut@reseau.eseo.fr> * Style Co-authored-by: Lysandre Debut <lysandre.debut@reseau.eseo.fr>	2022-10-27 16:03:11 -04:00
Sylvain Gugger	6c24443ff5	Safetensors tf (#19900 ) * Wip * Add safetensors support for TensorFlow * First tests * Add final test for now * Retrigger CI like this * Update src/transformers/modeling_tf_utils.py Co-authored-by: Lysandre Debut <lysandre.debut@reseau.eseo.fr> Co-authored-by: Lysandre Debut <lysandre.debut@reseau.eseo.fr>	2022-10-27 15:56:29 -04:00
Steven Liu	e4132952a1	Add GPT2 resources (#19879 ) * add resources for gpt2 * add pipeline icons and community resources	2022-10-27 11:34:00 -07:00
Steven Liu	d818dd3a41	Add BLOOM resources (#19881 ) * add bloom resources * add pipeline icon	2022-10-27 11:33:52 -07:00

1 2 3 4 5 ...

11210 Commits