transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-12 17:20:03 +06:00

Author	SHA1	Message	Date
Arthur	4864d08d3e	Add-support for commit description (#26704 ) * fix * update * revert * add dosctring * good to go * update * add a test	2023-10-26 12:37:09 +02:00
Arthur	15cd096288	Create SECURITY.md	2023-10-26 12:26:47 +02:00
Younes Belkada	fe2877ce21	Remove unneeded prints in modeling_gpt_neox.py (#27080 )	2023-10-26 11:55:31 +02:00
Younes Belkada	efba1a1744	Bump`flash_attn` version to `2.1` (#27079 ) * pin FA-2 to `2.1` * fix on modeling	2023-10-26 11:21:04 +02:00
Zach Mueller	90412401e6	Bring back `set_epoch` for Accelerate-based dataloaders (#26850 ) * Working tests! * Fix sampler * Fix * Update src/transformers/trainer.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Fix check * Clean --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2023-10-26 11:20:11 +02:00
dependabot[bot]	3c2692407d	Bump urllib3 from 1.26.17 to 1.26.18 in /examples/research_projects/lxmert (#26888 ) Bump urllib3 in /examples/research_projects/lxmert Bumps [urllib3](https://github.com/urllib3/urllib3) from 1.26.17 to 1.26.18. - [Release notes](https://github.com/urllib3/urllib3/releases) - [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst) - [Commits](https://github.com/urllib3/urllib3/compare/1.26.17...1.26.18) --- updated-dependencies: - dependency-name: urllib3 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2023-10-26 09:10:29 +02:00
dependabot[bot]	9c5240af14	Bump werkzeug from 2.2.3 to 3.0.1 in /examples/research_projects/decision_transformer (#27072 ) Bump werkzeug in /examples/research_projects/decision_transformer Bumps [werkzeug](https://github.com/pallets/werkzeug) from 2.2.3 to 3.0.1. - [Release notes](https://github.com/pallets/werkzeug/releases) - [Changelog](https://github.com/pallets/werkzeug/blob/main/CHANGES.rst) - [Commits](https://github.com/pallets/werkzeug/compare/2.2.3...3.0.1) --- updated-dependencies: - dependency-name: werkzeug dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>	2023-10-26 08:56:28 +02:00
corey hu	df2eebf1e7	Handle unsharded Llama2 model types in conversion script (#27069 ) Handle all unshared models types	2023-10-26 08:41:07 +02:00
Aarya Balwadkar	a2f55a65cd	Hindi translation of pipeline_tutorial.md (#26837 ) * hindi translation of pipeline_tutorial.md * Update pipeline_tutorial.md * Update build_documentation.yml * Update build_pr_documentation.yml * Updated build_documentation.yml --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2023-10-25 11:21:49 -07:00
Yeyang	ba5144f7a9	🌐 [i18n-ZH] Translate custom_models.md into Chinese (#27065 ) * docs(zh): translate custom_models.md * minor fix in customer_models Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2023-10-25 11:20:32 -07:00
Younes Belkada	c34c50cdc0	[`docs`] Add `MaskGenerationPipeline` in docs (#27063 ) * add `MaskGenerationPipeline` in docs * Update __init__.py * fix repo consistency and clarify docstring * add on check docstirngs * actually we do have a tf sam * oops	2023-10-25 19:31:36 +02:00
Akash Kundu	ba073ea9e3	[DOCS] minor fixes in README.md (#27048 ) minor fixes	2023-10-25 10:21:13 -07:00
Jing Hua	a64f8c1f87	[docstring] fix incorrect llama docstring: encoder -> decoder (#27071 ) fix incorrect docstring: encoder -> decoder	2023-10-25 18:09:04 +02:00
Nick Hill	0baa9246cb	Fix TypicalLogitsWarper tensor OOB indexing edge case (#26579 ) * Fix TypicalLogitsWarper tensor OOB indexing edge case This can be triggerd fairly quickly with low precision e.g. bfloat16 and typical_p = 0.99. * Shift threshold index by one * Use explicit named arg for clamp min	2023-10-25 11:36:43 +01:00
Younes Belkada	06e782da4e	[`core`] Refactor of `gradient_checkpointing` (#27020 ) * v1 * fix * remove `create_custom_forward` * fixup * fixup * add test and fix all failing GC tests * remove all remaining `create_custom_forward` methods * fix idefics bug * fixup * replace with `__call__` * add comment * quality	2023-10-25 12:16:15 +02:00
Arthur	9286f0ac39	Skip-test (#27062 ) * skip plbart test * nits * update	2023-10-25 10:47:33 +02:00
Tom Aarsen	6cbc1369a3	Fix RoPE config validation for FalconConfig + various config typos (#26929 ) * Resolve incorrect ValueError in RoPE config for Falcon * Add broken codeblock tag in Falcon Config * Fix typo: an float -> a float * Implement copy functionality for Fuyu and Persimmon for RoPE scaling validation * Make style	2023-10-24 18:37:09 +01:00
JB (Don)	a0fd34483f	Add a default decoder_attention_mask for EncoderDecoderModel during training (#26752 ) * Add a default decoder_attention_mask for EncoderDecoderModel during training Since we are already creating the default decoder_input_ids from the labels, we should also create a default decoder_attention_mask to go with it. * Fix test constant that relied on manual_seed() The test was changed to use a decoder_attention_mask that ignores padding instead (which is the default one created by BERT when attention_mask is None). * Create the decoder_attention_mask using decoder_input_ids instead of labels * Fix formatting in test	2023-10-24 18:26:16 +01:00
Maria Khalusova	9333bf0769	[docs] Performance docs refactor p.2 (#26791 ) * initial edits * improvements for clarity and flow * improvements for clarity and flow, removed the repetead section * removed two docs that had no content * Revert "removed two docs that had no content" This reverts commit `e98fa2fa0d`. * Apply suggestions from code review Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * feedback addressed * more feedback addressed * feedback addressed --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2023-10-24 13:10:06 -04:00
Patrick von Platen	13ef14e18e	Fix config silent copy in from_pretrained (#27043 ) * Fix config modeling utils * fix more * fix attn mask bug * Update src/transformers/modeling_utils.py	2023-10-24 19:05:37 +02:00
Alex McKinney	9da451713d	Device agnostic testing (#25870 ) * adds agnostic decorators and availability fns * renaming decorators and fixing imports * updating some representative example tests bloom, opt, and reformer for now * wip device agnostic functions * lru cache to device checking functions * adds `TRANSFORMERS_TEST_DEVICE_SPEC` if present, imports the target file and updates device to function mappings * comments `TRANSFORMERS_TEST_DEVICE_SPEC` code * extra checks on device name * `make style; make quality` * updates default functions for agnostic calls * applies suggestions from review * adds `is_torch_available` guard * Add spec file to docs, rename function dispatch names to backend_* * add backend import to docs example for spec file * change instances of to * Move register backend to before device check as per @statelesshz changes * make style * make opt test require fp16 to run --------- Co-authored-by: arsalanu <arsalanu@graphcore.ai> Co-authored-by: arsalanu <hzji210@gmail.com>	2023-10-24 16:49:26 +02:00
Marc Sun	41496b95da	Add fuyu device map (#26949 ) * add _no_split_modules * style * fix _no_split_modules * add doc	2023-10-24 09:10:23 -04:00
Leandro von Werra	b18e31407c	add info on TRL docs (#27024 ) * add info on TRL docs * add TRL link * tweak text * tweak text	2023-10-24 14:56:00 +02:00
amyeroberts	cb0c68069d	Safe import of rgb_to_id from FE modules (#27037 ) Safe import from FE modules	2023-10-24 13:40:16 +01:00
Arthur	7bde5d634f	[`TFxxxxForSequenceClassifciation`] Fix the eager mode after #25085 (#25751 ) * TODOS * Switch .shape -> shape_list --------- Co-authored-by: Matt <rocketknight1@gmail.com>	2023-10-24 13:33:05 +01:00
Michal Jamroz	e2d6d5ce57	Normalize only if needed (#26049 ) * Normalize only if needed * Update examples/pytorch/image-classification/run_image_classification.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * if else in one line * within block * one more place, sorry for mess * import order * Update examples/pytorch/image-classification/run_image_classification.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update examples/pytorch/image-classification/run_image_classification_no_trainer.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> --------- Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2023-10-24 13:32:03 +01:00
JP	576e2823a3	Add descriptive docstring to WhisperTimeStampLogitsProcessor (#25642 ) * adding in logit examples for Whisper processor * adding in updated logits processor for Whisper * adding in cleaned version of logits processor for Whisper * adding docstrings for whisper processor * making sure the formatting is correct * adding logits after doc builder * Update src/transformers/generation/logits_process.py Adding in suggested fix to the LogitProcessor description. Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com> * Update src/transformers/generation/logits_process.py Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com> * Update src/transformers/generation/logits_process.py Removing tip per suggestion. Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com> * Update src/transformers/generation/logits_process.py Removing redundant code per suggestion. Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com> * adding in revised version * adding in version with timestamp examples * Update src/transformers/generation/logits_process.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * enhanced paragraph on behavior of processor * fixing doc quality issue * removing the word poem from example * adding in updated docstring * adding in new version of file after doc-builder --------- Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com> Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2023-10-24 12:02:06 +02:00
Yih-Dar	fc142bd775	Add `default_to_square_for_size` to `CLIPImageProcessor` (#26965 ) * fix * fix * fix * fix * fix --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2023-10-24 11:08:17 +02:00
Xuehai Pan	cc7803c0a6	Register ModelOutput as supported torch pytree nodes (#26618 ) * Register ModelOutput as supported torch pytree nodes * Test ModelOutput as supported torch pytree nodes * Update type hints for pytree unflatten functions	2023-10-24 11:02:40 +02:00
fxmarty	ede051f1b8	Fix key dtype in GPTJ and CodeGen (#26836 ) * fix key dtype in gptj and codegen * delay the key cast to a later point * fix	2023-10-24 16:55:14 +09:00
Yeyang	32f799db0d	🌐 [i18n-ZH] Translate create_a_model.md into Chinese (#27026 ) docs(zh): translate create_a_model.md	2023-10-23 15:44:42 -07:00
Mert Yanık	25c022d7c5	Fix little typo (#27028 )	2023-10-23 15:36:42 -07:00
Pedro Gabriel Gengo Lourenço	f370bebdc3	Bugfix device map detr model (#26849 ) * Fixed replace_batch_norm when on meta device * lint fix * Adding coauthor Co-authored-by: Pi Esposito <piero.skywalker@gmail.com> * Removed tests * Remove unused deps * Try to fix copy issue * try fix copy one more time * Reverted import changes --------- Co-authored-by: Pi Esposito <piero.skywalker@gmail.com>	2023-10-23 14:34:27 -04:00
jiaqiw09	b0d1d7f71a	translate `preprocessing.md` to Chinese (#26955 ) * translate preprocessing.md to Chinese * update files fixing problems mentioned in review * update files fixing problems mentioned in review --------- Co-authored-by: jiaqiw <wangjiaqi50@huawei.com>	2023-10-23 10:36:24 -07:00
Yeyang	19ae0505ae	🌐 [i18n-ZH] Translate multilingual into Chinese (#26935 ) translate multilingual into Chinese Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2023-10-23 10:35:17 -07:00
Patrick von Platen	33f98cfded	Remove ambiguous `padding_mask` and instead use a 2D->4D Attn Mask Mapper (#26792 ) * [Attn Mask Converter] refactor attn mask * up * Apply suggestions from code review Co-authored-by: fxmarty <9808326+fxmarty@users.noreply.github.com> * improve * rename * better cache * renaming * improve more * improve * fix bug * finalize * make style & make fix-copies * correct more * start moving attention_mask * fix llama * improve falcon * up * improve more * improve more * Update src/transformers/models/owlv2/modeling_owlv2.py * make style * make style * rename to converter * Apply suggestions from code review --------- Co-authored-by: fxmarty <9808326+fxmarty@users.noreply.github.com>	2023-10-23 18:54:00 +02:00
jiaqiw09	f09a081d27	Translate `pipeline_tutorial.md` to chinese (#26954 ) * update translation of pipeline_tutorial and preprocessing(Version1.0) * update translation of pipeline_tutorial and preprocessing(Version2.0) * update translation docs * update to fix problems mentioned in review --------- Co-authored-by: jiaqiw <wangjiaqi50@huawei.com>	2023-10-23 08:58:00 -07:00
Matt	f7354a3bd6	Remove token_type_ids from default TF GPT-2 signature (#26962 ) Remove token_type_ids from default GPT-2 signature	2023-10-23 16:18:02 +01:00
Rafael Padilla	c0b5ad9473	small typos found (#26988 ) just very small typos found	2023-10-23 11:08:39 -03:00
Arthur	f9f27b0fc2	[`SeamlessM4T`] fix copies with NLLB MoE int8 (#27018 ) fix copies on newly merged model	2023-10-23 15:25:06 +02:00
Younes Belkada	244a53e0f6	[`NLLB-MoE`] Fix NLLB MoE 4bit inference (#27012 ) fix NLLB MoE 4bit	2023-10-23 14:54:22 +02:00
Yoach Lacombe	cb45f71c4d	Add Seamless M4T model (#25693 ) * first raw commit * still POC * tentative convert script * almost working speech encoder conversion scripts * intermediate code for encoder/decoders * add modeling code * first version of speech encoder * make style * add new adapter layer architecture * add adapter block * add first tentative config * add working speech encoder conversion * base model convert works now * make style * remove unnecessary classes * remove unecessary functions * add modeling code speech encoder * rework logics * forward pass of sub components work * add modeling codes * some config modifs and modeling code modifs * save WIP * new edits * same output speech encoder * correct attention mask * correct attention mask * fix generation * new generation logics * erase comments * make style * fix typo * add some descriptions * new state * clean imports * add tests * make style * make beam search and num_return_sequences>1 works * correct edge case issue * correct SeamlessM4TConformerSamePadLayer copied from * replace ACT2FN relu by nn.relu * remove unecessary return variable * move back a class * change name conformer_attention_mask ->conv_attention_mask * better nit code * add some Copied from statements * small nits * small nit in dict.get * rename t2u model -> conditionalgeneration * ongoing refactoring of structure * update models architecture * remove SeamlessM4TMultiModal classes * add tests * adapt tests * some non-working code for vocoder * add seamlessM4T vocoder * remove buggy line * fix some hifigan related bugs * remove hifigan specifc config * change * add WIP tokenization * add seamlessM4T working tokenzier * update tokenization * add tentative feature extractor * Update converting script * update working FE * refactor input_values -> input_features * update FE * changes in generation, tokenizer and modeling * make style and add t2u_decoder_input_ids * add intermediate outputs for ToSpeech models * add vocoder to speech models * update valueerror * update FE with languages * add vocoder convert * update config docstrings and names * update generation code and configuration * remove todos and update config.pad_token_id to generation_config.pad_token_id * move block vocoder * remove unecessary code and uniformize tospeech code * add feature extractor import * make style and fix some copies from * correct consistency + make fix-copies * add processor code * remove comments * add fast tokenizer support * correct pad_token_id in M4TModel * correct config * update tests and codes + make style * make some suggested correstion - correct comments and change naming * rename some attributes * rename some attributes * remove unecessary sequential * remove option to use dur predictor * nit * refactor hifigan * replace normalize_mean and normalize_var with do_normalize + save lang ids to generation config * add tests * change tgt_lang logic * update generation ToSpeech * add support import SeamlessM4TProcessor * fix generate * make tests * update integration tests, add option to only return text and update tokenizer fast * fix wrong function call * update import and convert script * update integration tests + update repo id * correct paths and add first test * update how new attention masks are computed * update tests * take first care of batching in vocoder code * add batching with the vocoder * add waveform lengths to model outputs * make style * add generate kwargs + forward kwargs of M4TModel * add docstrings forward methods * reformate docstrings * add docstrings t2u model * add another round of modeling docstrings + reformate speaker_id -> spkr_id * make style * fix check_repo * make style * add seamlessm4t to toctree * correct check_config_attributes * write config docstrings + some modifs * make style * add docstrings tokenizer * add docstrings to processor, fe and tokenizers * make style * write first version of model docs * fix FE + correct FE test * fix tokenizer + add correct integration tests * fix most tokenization tests * make style * correct most processor test * add generation tests and fix num_return_sequences > 1 * correct integration tests -still one left * make style * correct position embedding * change numbeams to 1 * refactor some modeling code and correct one test * make style * correct typo * refactor intermediate fnn * refactor feedforward conformer * make style * remove comments * make style * fix tokenizer tests * make style * correct processor tests * make style * correct S2TT integration * Apply suggestions from Sanchit code review Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com> * correct typo * replace torch.nn->nn + make style * change Output naming (waveforms -> waveform) and ordering * nit renaming and formating * remove return None when not necessary * refactor SeamlessM4TConformerFeedForward * nit typo * remove almost copied from comments * add a copied from comment and remove an unecessary dropout * remove inputs_embeds from speechencoder * remove backward compatibiliy function * reformate class docstrings for a few components * remove unecessary methods * split over 2 lines smthg hard to read * make style * replace two steps offset by one step as suggested * nice typo * move warnings * remove useless lines from processor * make generation non-standard test more robusts * remove torch.inference_mode from tests * split integration tests * enrich md * rename control_symbol_vocoder_offset->vocoder_offset * clean convert file * remove tgt_lang and src_lang from FE * change generate docstring of ToText models * update generate docstring of tospeech models * unify how to deal withtext_decoder_input_ids * add default spkr_id * unify tgt_lang for t2u_model * simplify tgt_lang verification * remove a todo * change config docstring * make style * simplify t2u_tgt_lang_id * make style * enrich/correct comments * enrich .md * correct typo in docstrings * add torchaudio dependency * update tokenizer * make style and fix copies * modify SeamlessM4TConverter with new tokenizer behaviour * make style * correct small typo docs * fix import * update docs and add requirement to tests * add convert_fairseq2_to_hf in utils/not_doctested.txt * update FE * fix imports and make style * remove torchaudio in FE test * add seamless_m4t.md to utils/not_doctested.txt * nits and change the way docstring dataset is loaded * move checkpoints from ylacombe/ to facebook/ orga * refactor warning/error to be in the 119 line width limit * round overly precised floats * add stereo audio behaviour * refactor .md and make style * enrich docs with more precised architecture description * readd undocumented models * make fix-copies * apply some suggestions * Apply suggestions from code review Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com> Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * correct bug from previous commit * refactor a parameter allowing to clean the code + some small nits * clean tokenizer * make style and fix * make style * clean tokenizers arguments * add precisions for some tests * move docs from not_tested to slow * modify tokenizer according to last comments * add copied from statements in tests * correct convert script * correct parameter docstring style * correct tokenization * correct multi gpus * make style * clean modeling code * make style * add copied from statements * add copied statements * add support with ASR pipeline * remove file added inadvertently * fix docstrings seamlessM4TModel * add seamlessM4TConfig to OBJECTS_TO_IGNORE due of unconventional markdown * add seamlessm4t to assisted generation ignored models --------- Co-authored-by: Sanchit Gandhi <93869735+sanchit-gandhi@users.noreply.github.com> Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2023-10-23 14:49:48 +02:00
Younes Belkada	50d0cf4f6b	Change default `max_shard_size` to smaller value (#26942 ) * Update modeling_utils.py * fixup * let's change it to 5GB * fix	2023-10-23 14:25:48 +02:00
Omar Sanseviero	d33d313192	Nits in Llama2 docstring (#26996 ) Update llama2.md	2023-10-23 14:19:59 +02:00
Arthur	ef978d0a7b	skip two tests (#27013 ) * skip two tests * skip torch as well * fixup	2023-10-23 12:52:05 +02:00
Gema Parreño	45425660d0	python falcon doc-string example typo (#26995 ) git python falcon typo	2023-10-23 12:51:35 +02:00
Lysandre Debut	700329493d	Limit to inferior fsspec version (#27010 ) Pin fsspec	2023-10-23 12:34:21 +02:00
YQ	f71c9ccf59	fix logit-to-multi-hot conversion in example (#26936 ) * fix logit to multi-hot converstion * add comments * typo	2023-10-23 12:33:05 +02:00
Akhil	093848d3cc	Added Telugu [te] translations (#26828 ) * Create index.md * Create _toctree.yml * Updated index.md in telugu * Update _toctree.yml * Create quicktour.md * Update quicktour.md * Create index.md * Update quicktour.md * Update docs/source/te/quicktour.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Delete docs/source/hi/index.md * Update docs/source/te/quicktour.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/te/quicktour.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/te/quicktour.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/te/quicktour.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/te/quicktour.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/te/quicktour.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/te/quicktour.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/te/quicktour.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update build_documentation.yml Added telugu [te] * Update build_pr_documentation.yml Added Telugu [te] * Update _toctree.yml --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2023-10-20 15:27:55 -07:00
Biswa Baibhab Subudhi	224794b011	Update README_hd.md (#26872 ) * Update README_hd.md - Fixed broken links I hope this small contribution adds value to this project. * Update README_hd.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2023-10-20 14:23:41 -07:00

... 14 15 16 17 18 ...

15053 Commits