transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-31 02:02:21 +06:00

Author	SHA1	Message	Date
regisss	979fccc90f	Enable BLIP for auto VQA (#29499 ) * Enable BLIP for auto VQA * Make style * Add VQA to BLIP pipeline tests	2024-03-07 10:28:01 +01:00
Park Jun	d45f47ab7f	Fix: Disable torch.autocast in RotaryEmbedding of Gemma and LLaMa for MPS device (#29439 ) * Fix: Disable torch.autocast in RotaryEmbedding of Gemma and LLaMa for MPS devices * Update src/transformers/models/gemma/modeling_gemma.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update llama ang gemma rope use cpu in mps device --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2024-03-07 00:57:22 +01:00
Glen Taggart	2a939f20ff	Substantially reduce memory usage in _update_causal_mask for large batches by using .expand instead of .repeat [needs tests+sanity check] (#29413 ) * try to fix gemma mem use * fix: handle attention mask dim==2 case * remove logits=logits.float() * clean up + add llama * apply formatting * readability edit: swap order of items being multiplied * revert change unrelated to PR * revert black autoformat * switch to one .to * Accept style edits Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2024-03-07 00:56:25 +01:00
Alvaro Bartolome	965cf67769	Fix `TextGenerationPipeline.__call__` docstring (#29491 )	2024-03-06 09:03:55 -08:00
Moshe Berchansky	19fb1e22d2	added the max_matching_ngram_size to GenerationConfig (#29131 ) * added the max_matching_ngram_size parameter into the GenerationConfig, for the PromptLookupCandidateGenerator * switched back to keyword arguments * added PromptLookupCandidateGenerator docstring for its parameters * ruff reformat * Update src/transformers/generation/configuration_utils.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> --------- Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com> Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2024-03-06 15:06:45 +00:00
Joao Gante	ddb4fda3cb	Generate: torch.compile-ready generation config preparation (#29443 )	2024-03-06 14:28:45 +00:00
Zach Mueller	9322576e2f	Fix test failure on DeepSpeed (#29444 ) * Fix test failure * use item	2024-03-06 07:11:53 -05:00
Ofir Zafrir	0a5b0516f8	Avoid dummy token in PLD to optimize performance (#29445 )	2024-03-06 11:19:47 +00:00
Joao Gante	700d48fb2d	Generate: get generation mode from the generation config instance 🧼 (#29441 )	2024-03-06 11:18:35 +00:00
Joao Gante	41f7b7ae4b	Generate: add tests for caches with `pad_to_multiple_of` (#29462 )	2024-03-06 10:57:04 +00:00
Matthew Hoffman	2890116ab7	Fix TrainingArguments regression with torch <2.0.0 for dataloader_prefetch_factor (#29447 ) * Fix TrainingArguments regression with torch <2.0.0 for dataloader_prefetch_factor dataloader_prefetch_factor was added to TrainingArguments in #28498 with the default value None, but versions of torch<2.0.0 do not accept None and will raise an error if num_workers == 0 and prefetch_factor != 2 * Add is_torch_available() check * Use is_torch_greater_or_equal_than_2_0 add back check for dataloader_prefetch_factor	2024-03-06 09:44:08 +00:00
Younes Belkada	b27aa206dd	[`docs`] Add starcoder2 docs (#29454 ) * add accelerate docs * Apply suggestions from code review Co-authored-by: Loubna Ben Allal <44069155+loubnabnl@users.noreply.github.com> * Update starcoder2.md * add correct generation --------- Co-authored-by: Loubna Ben Allal <44069155+loubnabnl@users.noreply.github.com>	2024-03-06 06:58:37 +01:00
Younes Belkada	2a002d073a	[`Docs` / `Awq`] Add docs on exllamav2 + AWQ (#29474 ) * add docs on exllamav2 + AWQ * Update docs/source/en/quantization.md	2024-03-06 06:30:47 +01:00
Fanli Lin	00bf44270f	[FIX] `offload_weight()` takes from 3 to 4 positional arguments but 5 were given (#29457 ) * use require_torch_gpu * enable on XPU * fix	2024-03-06 03:58:42 +01:00
AI4Harmony	7b01579f73	🌐 [i18n-KO] Translated generation_strategies.md to Korean (#29086 ) * Update ko _toctree.yml * Create ko: generation_strategies.md * Apply suggestions from code review Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Apply suggestions from code review Co-authored-by: Jungnerd <46880056+jungnerd@users.noreply.github.com> * Apply suggestions from code review Co-authored-by: Jungnerd <46880056+jungnerd@users.noreply.github.com> --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> Co-authored-by: Jungnerd <46880056+jungnerd@users.noreply.github.com>	2024-03-05 15:47:33 -08:00
Michael	638c423c89	[i18n-zh] Translate add_new_pipeline.md into Chinese (#29432 ) * [i18n-zh] Translate add_new_pipeline.md into Chinese * apply suggestions from Fan-Lin	2024-03-05 09:19:00 -08:00
Lysandre Debut	a69cbf4e64	Automatic safetensors conversion when lacking these files (#29390 ) * Automatic safetensors conversion when lacking these files * Remove debug * Thread name * Typo * Ensure that raises do not affect the main thread	2024-03-05 13:37:55 +01:00
Logan Adams	9c5e560924	Update pytest `import_path` location (#29154 ) * Update to pull function from proper lib * Fix ruff formatting error * Remove accidently added file	2024-03-05 12:23:34 +00:00
AleksanderWWW	8f3f8e6766	Fix bug with passing capture_* args to neptune callback (#29041 ) * Fix bug with passing capture_* args to neptune callback * ruff happy? * instantiate (frozen)set only once * code review * code review 2 * ruff happy? * code review	2024-03-05 11:54:00 +00:00
Arthur	fb1c62e973	[`Add Mamba`] Adds support for the `Mamba` models (#28094 ) * initial-commit * start cleaning * small nits * small nits * current updates * add kernels * small refactoring little step * add comments * styling * nit * nits * Style * Small changes * Push dummy mambda simple slow * nit * Use original names * Use original names and remove norm * Updates for inference params * Style nd updates * nits * Match logits * Add a test * Add expected generated text * nits doc, imports and styling * style * oups * dont install kernels, invite users to install the required kernels * let use use the original packages * styling * nits * fix some copieds * update doc * fix-copies * styling done * nits * fix import check * run but wrong cuda ress * mamba CUDA works :) * fix the fast path * config naming nits * conversion script is not required at this stage * finish fixing the fast path: generation make sense now! * nit * Let's start working on the CIs * style * better style * more nits * test nit * quick fix for now * nits * nit * nit * nit * nits * update test rest * fixup * update test * nit * some fixes * nits * update test values * fix styling * nit * support peft * integrations tests require torchg * also add slow markers * styling * chose forward wisely * nits * update tests * fix gradient checkpointing * fixup * nit * fix doc * check copies * fix the docstring * fix some more tests * style * fix beam search * add init schene * update * nit * fix * fixup the doc * fix the doc * fixup * tentative update but slow is no longer good * nit * should we always use float32? * nits * revert wrong changes * res in float32 * cleanup * skip fmt for now * update generation values * update test values running original model * fixup * update tests + rename inference_params to cache_params + make sure training does not use cache_params * small nits * more nits * fix final CIs * style * nit doc * I hope final doc nits * nit * 🫠 * final touch! * fix torch import * Apply suggestions from code review Co-authored-by: Lysandre Debut <hi@lysand.re> * Apply suggestions from code review * fix fix and fix * fix base model prefix! * nit * Update src/transformers/models/mamba/__init__.py * Update docs/source/en/model_doc/mamba.md Co-authored-by: Lysandre Debut <hi@lysand.re> * nit --------- Co-authored-by: Lysandre Debut <hi@lysand.re>	2024-03-05 20:01:06 +09:00
Joao Gante	87a0783dde	Generate: inner decoding methods are no longer public (#29437 )	2024-03-05 10:27:36 +00:00
Arthur	4d892b7297	[`Udop imports`] Processor tests were not run. (#29456 ) * fix udop imports * sort imports	2024-03-05 11:01:08 +01:00
Arthur	57d007b912	Revert-commit `0d52f9f582` (#29455 ) * style * revert with RP * nit * exact revert	2024-03-05 10:39:42 +01:00
Arthur Zucker	0d52f9f582	more fix	2024-03-05 18:27:25 +09:00
Arthur	132852203a	[`UdopTokenizer`] Fix post merge imports (#29451 ) * update * ... * nits * arf * 🧼 * beat the last guy * style everyone	2024-03-05 09:42:52 +01:00
Fanli Lin	fa7f3cf336	[tests] enable test_pipeline_accelerate_top_p on XPU (#29309 ) * use torch_device * Update tests/pipelines/test_pipelines_text_generation.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * fix style --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2024-03-05 09:16:05 +01:00
Joshua Lochner	ebccb09169	[docs] Update starcoder2 paper link (#29418 ) Update starcoder2 paper link	2024-03-05 08:57:33 +01:00
Raushan Turganbay	bd891aed01	Fix max length for BLIP generation (#29296 ) * fix mal_length for blip * update also min length * fixes * add a comment * Update src/transformers/models/instructblip/modeling_instructblip.py Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com> * Update src/transformers/models/blip_2/modeling_blip_2.py Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com> * make fixup * fix length when user passed * remove else * remove brackets --------- Co-authored-by: Joao Gante <joaofranciscocardosogante@gmail.com>	2024-03-05 08:18:22 +01:00
Ilyas Moutawwakil	4fc708f98c	Exllama kernels support for AWQ models (#28634 ) * added exllama kernels support for awq models * doc * style * Update src/transformers/modeling_utils.py Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com> * refactor * moved exllama post init to after device dispatching * bump autoawq version * added exllama test * style * configurable exllama kernels * copy exllama_config from gptq * moved exllama version check to post init * moved to quantization dockerfile --------- Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>	2024-03-05 03:22:48 +01:00
Younes Belkada	81c8191b46	FIX [`Generation`] Fix some issues when running the MaxLength criteria on CPU (#29317 ) fix the bitwise or issue	2024-03-05 02:29:19 +01:00
njackman-2344	e947683294	[Docs] Spanish Translation -Torchscript md & Trainer md (#29310 ) * torchscript and trainer md es translation * corrected md es files and even corrected spelling in en md * made es corrections to trainer.md * deleted entrenamiento... title on yml * placed entrenamiento in right place	2024-03-04 13:57:51 -08:00
NielsRogge	836921fdeb	Add UDOP (#22940 ) * First draft * More improvements * More improvements * More fixes * Fix copies * More improvements * More fixes * More improvements * Convert checkpoint * More improvements, set up tests * Fix more tests * Add UdopModel * More improvements * Fix equivalence test * More fixes * Redesign model * Extend conversion script * Use real inputs for conversion script * Add image processor * Improve conversion script * Add UdopTokenizer * Add fast tokenizer * Add converter * Update README's * Add processor * Add fully fledged tokenizer * Add fast tokenizer * Use processor in conversion script * Add tokenizer tests * Fix one more test * Fix more tests * Fix tokenizer tests * Enable fast tokenizer tests * Fix more tests * Fix additional_special_tokens of fast tokenizer * Fix tokenizer tests * Fix more tests * Fix equivalence test * Rename image to pixel_values * Rename seg_data to bbox * More renamings * Remove vis_special_token * More improvements * Add docs * Fix copied from * Update slow tokenizer * Update fast tokenizer design * Make text input optional * Add first draft of processor tests * Fix more processor tests * Fix decoder_start_token_id * Fix test_initialization * Add integration test * More improvements * Improve processor, add test * Add more copied from * Add more copied from * Add more copied from * Add more copied from * Remove print statement * Update README and auto mapping * Delete files * Delete another file * Remove code * Fix test * Fix docs * Remove asserts * Add doc tests * Include UDOP in exotic model tests * Add expected tesseract decodings * Add sentencepiece * Use same design as T5 * Add UdopEncoderModel * Add UdopEncoderModel to tests * More fixes * Fix fast tokenizer * Fix one more test * Remove parallelisable attribute * Fix copies * Remove legacy file * Copy from T5Tokenizer * Fix rebase * More fixes, copy from T5 * More fixes * Fix init * Use ArthurZ/udop for tests * Make all model tests pass * Remove UdopForConditionalGeneration from auto mapping * Fix more tests * fixups * more fixups * fix the tokenizers * remove un-necessary changes * nits * nits * replace truncate_sequences_boxes with truncate_sequences for fix-copies * nit current path * add a test for input ids * ids that we should get taken from `c9f7a32f57` * nits converting * nits * apply ruff * nits * nits * style * fix slow order of addition * fix udop fast range as well * fixup * nits * Add docstrings * Fix gradient checkpointing * Update code examples * Skip tests * Update integration test * Address comment * Make fixup * Remove extra ids from tokenizer * Skip test * Apply suggestions from code review Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update year * Address comment * Address more comments * Address comments * Add copied from * Update CI * Rename script * Update model id * Add AddedToken, skip tests * Update CI * Fix doc tests * Do not use Tesseract for the doc tests * Remove kwargs * Add original inputs * Update casting * Fix doc test * Update question * Update question * Use LayoutLMv3ImageProcessor * Update organization * Improve docs * Update forward signature * Make images optional * Remove deprecated device argument * Add comment, add add_prefix_space * More improvements * Remove kwargs --------- Co-authored-by: ArthurZucker <arthur.zucker@gmail.com> Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2024-03-04 18:49:02 +01:00
Donggeun Yu	ed74d97871	DeformableDETR support bfloat16 (#29232 ) * Update ms_deform_attn_cuda.cu * Update ms_deform_attn_cuda.cuh * Update modeling_deformable_detr.py * Update src/transformers/models/deformable_detr/modeling_deformable_detr.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update modeling_deformable_detr.py * python utils/check_copies.py --fix_and_overwrite * Fix dtype missmatch error * Update test_modeling_deformable_detr.py * Update test_modeling_deformable_detr.py * Update modeling_deformable_detr.py * Update modeling_deformable_detr.py * Support DeformableDETR with bfloat16 * Add test code * Use AT_DISPATCH_FLOATING_TYPES_AND2 Use AT_DISPATCH_FLOATING_TYPES_AND2 * Update tests/models/deformable_detr/test_modeling_deformable_detr.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update tests/models/deformable_detr/test_modeling_deformable_detr.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Fix not found require_torch_bf16 function --------- Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2024-03-04 14:18:09 +00:00
Yoach Lacombe	bcd23a54f1	Avoid edge case in audio utils (#28836 )	2024-03-04 13:24:40 +00:00
Sven Schultze	7941769e55	Fix grad_norm unserializable tensor log failure (#29212 ) * Fix grad_norm unserializable tensor log failure * Fix origin of grad_norm logs to be in deepspeed get_global_grad_norm()	2024-03-04 13:12:35 +00:00
Zach Mueller	1681a6d452	🚨 Fully revert atomic checkpointing 🚨 (#29370 ) Fully revert atomic checkpointing	2024-03-04 06:17:42 -05:00
Nick DeGroot	8ef9862864	Fix OneFormer `post_process_instance_segmentation` for panoptic tasks (#29304 ) * 🐛 Fix oneformer instance post processing when using panoptic task type * ✅ Add unit test for oneformer instance post processing panoptic bug --------- Co-authored-by: Nick DeGroot <1966472+nickthegroot@users.noreply.github.com>	2024-03-04 11:04:49 +00:00
Sean (Seok-Won) Yi	81220cba61	Fix: Fixed the previous tracking URI setting logic to prevent clashes with original MLflow code. (#29096 ) * Changed logic for setting the tracking URI. The previous code was calling the `mlflow.set_tracking_uri` function regardless of whether or not the environment variable `MLFLOW_TRACKING_URI` is even set. This led to clashes with the original MLflow implementation and therefore the logic was changed to only calling the function when the environment variable is explicitly set. * Check if tracking URI has already been set. The previous code did not consider the possibility that the tracking URI may already be set elsewhere and was therefore (erroneously) overriding previously set tracking URIs using the environment variable. * Removed redundant parentheses. Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Fix docstring to reflect library convention properly. Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Fix docstring to reflect library convention properly. "Unset by default" is the correct expression rather than "Default to `None`." Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> --------- Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2024-03-04 10:53:58 +00:00
NielsRogge	5e4b69dc12	Convert SlimSAM checkpoints (#28379 ) * First commit * Improve conversion script * Convert more checkpoints * Update src/transformers/models/sam/convert_sam_original_to_hf_format.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Rename file * More updates * Update docstring * Update script --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2024-03-04 11:51:16 +01:00
Traun Leyden	c38a12270a	Workaround for #27758 to avoid ZeroDivisionError (#28756 )	2024-03-04 10:23:40 +01:00
Y4hL	704b3f74f9	Add mlx support to BatchEncoding.convert_to_tensors (#29406 ) * Add mlx support * Fix import order and use def instead of lambda * Another fix for ruff format :) * Add detecting mlx from repr, add is_mlx_array	2024-03-04 10:19:13 +01:00
Siming Dai	39ef3fb248	[Mixtral] Fixes attention masking in the loss (#29363 ) Fix mixtral load balancing loss Co-authored-by: dingkunbo <dingkunbo@baidu.com>	2024-03-04 09:08:56 +01:00
Poedator	38953a75c1	update path to hub files in the error message (#29369 ) update path to hub files need to add `tree/` to path to files at HF hub. see example path: `https://huggingface.co/meta-llama/Llama-2-7b-hf/tree/main`	2024-03-04 08:26:01 +01:00
Fanli Lin	aade711d1e	[tests] enable automatic speech recognition pipeline tests on XPU (#29308 ) * use require_torch_gpu * enable on XPU	2024-03-04 08:24:38 +01:00
David Valente	831bc25d8f	Correct zero division error in inverse sqrt scheduler (#28982 ) * Correct zero division error in inverse sqrt scheduler * default timescale to 10_000	2024-03-01 17:04:40 +00:00
Zach Mueller	1a7c117df9	Fix deprecated arg issue (#29372 ) * Fix deprecated arg issue * Trainer check too * Check for dict or dataclass * Simplify, make config always AcceleratorConfig * Upstream to Trainer	2024-03-01 12:00:29 -05:00
Marc Sun	cec773345a	Fix llama + gemma accelete tests (#29380 )	2024-03-01 10:32:36 -05:00
Jingya HUANG	15f8296a9b	Support subfolder with `AutoProcessor` (#29169 ) enable subfolder	2024-03-01 10:29:21 +00:00
amyeroberts	f1b1379f37	[`YOLOS`] Fix - return padded annotations (#29300 ) * Fix yolos processing * Add back slow marker - protects for pycocotools in slow * Slow decorator goes above copied from header	2024-03-01 09:42:13 +00:00
Sanchit Gandhi	0a0a279e99	🚨🚨[Whisper Tok] Update integration test (#29368 ) * [Whisper Tok] Update integration test * make style	2024-03-01 09:22:31 +00:00

1 2 3 4 5 ...

15287 Commits