transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-31 02:02:21 +06:00

Author	SHA1	Message	Date
Yih-Dar	73893df864	Fix `Owlv2ModelIntegrationTest::test_inference_object_detection` (#27793 ) * fix * fix --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2023-12-04 09:45:22 +01:00
Yih-Dar	5a551df92b	Fix `TvpModelIntegrationTests` (#27792 ) * fix * fix * fix --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2023-12-04 09:40:42 +01:00
Arthur	c0b9db0914	[`ModelOnTheFlyConversionTester`] Mark as slow for now (#27823 ) * mark test as slow for now * style	2023-12-04 08:33:15 +01:00
Ilya	269078a7eb	Add `persistent_workers` parameter to `TrainingArguments` (#27189 ) added param Co-authored-by: Ilya Fedorov <ilyaf@nvidia.com>	2023-12-04 07:43:32 +01:00
Noah Siegel	a2b1e1df49	Fix typo in max_length deprecation warnings (#27788 )	2023-12-04 07:41:50 +01:00
NielsRogge	7edf8bfafd	Improve forward signature test (#27729 ) * First draft * Extend test_forward_signature * Update tests/test_modeling_common.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Revert suggestion --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2023-12-04 07:38:22 +01:00
Roy Hvaara	bcd0a91a01	[JAX] Replace uses of jax.devices("cpu") with jax.local_devices(backend="cpu") (#27593 ) An upcoming change to JAX will include non-local (addressable) CPU devices in jax.devices() when JAX is used multicontroller-style, where there are multiple Python processes. This change preserves the current behavior by replacing uses of jax.devices("cpu"), which previously only returned local devices, with jax.local_devices("cpu"), which will return local devices both now and in the future. This change is always safe (i.e., it should always preserve the previous behavior), but it may sometimes be unnecessary if code is never used in a multicontroller setting. Co-authored-by: Peter Hawkins <phawkins@google.com>	2023-12-04 07:36:29 +01:00
Sanchit Gandhi	2c658b5a42	[MusicGen] Fix audio channel attribute (#27440 ) [MusicGen] Fix mono logit test	2023-12-01 17:10:03 +00:00
Marc Sun	abd4cbd775	Better error message for bitsandbytes import (#27764 ) * better error message * fix logic * fix log	2023-12-01 11:59:14 -05:00
Nicolas Patry	7b6324e18e	Make using safetensors files automated. (#27571 ) * [WIP] Make using safetensors files automated. If `use_safetensors=True` is used, and it doesn't exist: - Don't crash just yet - Lookup for an open PR containing it. - If yes, use that instead - If not, touch the space to convert, wait for conversion to be finished and the PR to be opened - Use that new PR - Profit. * Remove the token. * [Auto Safetensors] Websocket -> SSE (#27656) * Websocket -> SSE * Support sharded + tests +cleanup a * env var * Apply suggestions from code review * Thanks Simon * Thanks Wauplin Co-authored-by: Wauplin <lucainp@gmail.com> * Cleanup * Update tests * Tests should pass * Apply to other tests * Extend extension * relax requirement on latest hfh * Revert * Correct private handling & debug statements * Skip gated repos as of now * Address review comments Co-authored-by: ArthurZucker <arthur.zucker@gmail.com> --------- Co-authored-by: Lysandre Debut <hi@lysand.re> Co-authored-by: Lysandre <lysandre@huggingface.co> Co-authored-by: Wauplin <lucainp@gmail.com> Co-authored-by: Lysandre <lysandre.debut@reseau.eseo.fr> Co-authored-by: ArthurZucker <arthur.zucker@gmail.com>	2023-12-01 15:51:10 +01:00
Wesley Gifford	95900916ab	Fixes for PatchTST Config (#27777 ) * Remove config reference and pass num_patches for PatchTSTforPrediction * ensure return_dict is properly set --------- Co-authored-by: Wesley M. Gifford <wmgifford@us.ibm.com>	2023-12-01 14:57:50 +01:00
Nolwenn Bernard	cf62539a29	[i18n-fr] Translate installation to French (#27657 ) * partial traduction of installation * Finish translation of installation * Update installation.mdx * Rename installation.mdx to installation.md * Typos * Update docs/source/fr/installation.md Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update docs/source/fr/installation.md Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update docs/source/fr/installation.md Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update docs/source/fr/installation.md Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update docs/source/fr/installation.md Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update docs/source/fr/installation.md Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update docs/source/fr/installation.md Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update docs/source/fr/installation.md Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update docs/source/fr/installation.md Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update docs/source/fr/installation.md Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Address review comments --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2023-12-01 14:00:07 +01:00
Joshua Lochner	0ad4e7e6da	[SeamlessM4Tv2] Fix links in README (#27782 ) Fix typo in README	2023-12-01 10:39:33 +01:00
Liangliang-Ma	9ddbb696d2	Fix unsupported setting of self._n_gpu in training_args on XPU devices (#27716 ) change xpu _n_gpu = 1	2023-12-01 10:34:15 +01:00
Yoach Lacombe	29f1aee3b6	Add SeamlessM4T v2 (#27779 ) * add working convertion script * first non-working version of modeling code * update modeling code (working) * make style * make fix-copies * add config docstrings * add config to ignore docstrings formatage due to unconventional markdown * fix copies * fix generation num_return_sequences * enrich docs * add and fix tests beside integration tests * update integration tests * update repo id * add tie weights and make style * correct naming in .md * fix imports and so on * correct docstrings * fix fp16 speech forward * fix speechencoder attention * make style * fix copied from * rename SeamlessM4Tv2-v2 to SeamlessM4Tv2 * Apply suggestions on configuration Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * remove useless public models * fix private models + better naming for T2U models * clean speech encoder relative position embeddings * refactor chunk attention * add docstrings to chunk attention method * improve naming and docstrings * rename some attention variables + add temperature sampling in T2U model * rename DOCSTRINGS variable names * make style + remove 2 useless config parameters * enrich model card * remove any attention_head reference + fix temperature in T2U * new fmt and make style * Apply suggestions from code review Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * rename spkr_id->speaker_id and change docstrings of get_char_input_ids * simplify v2attention * make style * Update seamless_m4t_v2.md * update code and tests with last update * update repo ids * fill article name, abstract andauthors * update not_doctested and slow_doc tests --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2023-11-30 20:24:43 +01:00
Joao Gante	510270af34	Generate: `GenerationConfig` throws an exception when `generate` args are passed (#27757 )	2023-11-30 14:16:31 +00:00
Dave Berenbaum	fe41647afc	uses dvclive_test mode in examples/pytorch/test_accelerate_examples.py (#27763 )	2023-11-30 14:52:03 +01:00
Yih-Dar	62ab32b299	Remove `check_runner_status.yml` (#27767 ) fix Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2023-11-30 10:17:25 +01:00
Kevin Hu	083e36923a	Fix precision errors from casting rotary parameters to FP16 with AMP (#27700 ) * Update modeling_llama.py * Update modeling_open_llama.py * Update modeling_gpt_neox.py * Update modeling_mistral.py * Update modeling_persimmon.py * Update modeling_phi.py * Update modeling_falcon.py * Update modeling_gpt_neox_japanese.py	2023-11-29 16:30:49 +01:00
Kashif Rasul	af8acc4760	[Time series] Add patchtst (#27581 ) * add distribution head to forecasting * formatting * Add generate function for forecasting * Add generate function to prediction task * formatting * use argsort * add past_observed_mask ordering * fix arguments * docs * add back test_model_outputs_equivalence test * formatting * cleanup * formatting * use ACT2CLS * formatting * fix add_start_docstrings decorator * add distribution head and generate function to regression task add distribution head and generate function to regression task. Also made add PatchTSTForForecastingOutput, PatchTSTForRegressionOutput. * add distribution head and generate function to regression task add distribution head and generate function to regression task. Also made add PatchTSTForForecastingOutput, PatchTSTForRegressionOutput. * fix typos * add forecast_masking * fixed tests * use set_seed * fix doc test * formatting * Update docs/source/en/model_doc/patchtst.md Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com> * better var names * rename PatchTSTTranspose * fix argument names and docs string * remove compute_num_patches and unused class * remove assert * renamed to PatchTSTMasking * use num_labels for classification * use num_labels * use default num_labels from super class * move model_type after docstring * renamed PatchTSTForMaskPretraining * bs -> batch_size * more review fixes * use hidden_state * rename encoder layer and block class * remove commented seed_number * edit docstring * Add docstring * formatting * use past_observed_mask * doc suggestion * make fix-copies * use Args: * add docstring * add docstring * change some variable names and add PatchTST before some class names * formatting * fix argument types * fix tests * change x variable to patch_input * format * formatting * fix-copies * Update tests/models/patchtst/test_modeling_patchtst.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * move loss to forward * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * formatting * fix a bug when pre_norm is set to True * output_hidden_states is set to False as default * set pre_norm=True as default * format docstring * format * output_hidden_states is None by default * add missing docs * better var names * docstring: remove default to False in output_hidden_states * change labels name to target_values in regression task * format * fix tests * change to forecast_mask_ratios and random_mask_ratio * change mask names * change future_values to target_values param in the prediction class * remove nn.Sequential and make PatchTSTBatchNorm class * black * fix argument name for prediction * add output_attentions option * add output_attentions to PatchTSTEncoder * formatting * Add attention output option to all classes * Remove PatchTSTEncoderBlock * create PatchTSTEmbedding class * use config in PatchTSTPatchify * Use config in PatchTSTMasking class * add channel_attn_weights * Add PatchTSTScaler class * add output_attentions arg to test function * format * Update doc with image patchtst.md * fix-copies * rename Forecast <-> Prediction * change name of a few parameters to match with PatchTSMixer. * Remove ForForecasting class to match with other time series models. make style * Remove PatchTSTForForecasting in the test * remove PatchTSTForForecastingOutput class * change test_forecast_head to test_prediction_head * style * fix docs * fix tests * change num_labels to num_targets * Remove PatchTSTTranspose * remove arguments in PatchTSTMeanScaler * remove arguments in PatchTSTStdScaler * add config as an argument to all the scaler classes * reformat * Add norm_eps for batchnorm and layernorm * reformat. * reformat * edit docstring * update docstring * change variable name pooling to pooling_type * fix output_hidden_states as tuple * fix bug when calling PatchTSTBatchNorm * change stride to patch_stride * create PatchTSTPositionalEncoding class and restructure the PatchTSTEncoder * formatting * initialize scalers with configs * edit output_hidden_states * style * fix forecast_mask_patches doc string * doc improvements * move summary to the start * typo * fix docstring * turn off masking when using prediction, regression, classification * return scaled output * adjust output when using distribution head * remove _num_patches function in the config * get config.num_patches from patchifier init * add output_attentions docstring, remove tuple in output_hidden_states * change SamplePatchTSTPredictionOutput and SamplePatchTSTRegressionOutput to SamplePatchTSTOutput * remove print("model_class: ", model_class) * change encoder_attention_heads to num_attention_heads * change norm to norm_layer * change encoder_layers to num_hidden_layers * change shared_embedding to share_embedding, shared_projection to share_projection * add output_attentions * more robust check of norm_type * change dropout_path to path_dropout * edit docstring * remove positional_encoding function and add _init_pe in PatchTSTPositionalEncoding * edit shape of cls_token and initialize it * add a check on the num_input_channels. * edit head_dim in the Prediction class to allow the use of cls_token * remove some positional_encoding_type options, remove learn_pe arg, initalize pe * change Exception to ValueError * format * norm_type is "batchnorm" * make style * change cls_token shape * Change forecast_mask_patches to num_mask_patches. Remove forecast_mask_ratios. * Bring PatchTSTClassificationHead on top of PatchTSTForClassification * change encoder_ffn_dim to ffn_dim and edit the docstring. * update variable names to match with the config * add generation tests * change num_mask_patches to num_forecast_mask_patches * Add examples explaining the use of these models * make style * Revert "Revert "[time series] Add PatchTST (#25927)" (#27486)" This reverts commit `78f6ed6c70`. * make style * fix default std scaler's minimum_scale * fix docstring * close code blocks * Update docs/source/en/model_doc/patchtst.md Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update tests/models/patchtst/test_modeling_patchtst.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update src/transformers/models/patchtst/configuration_patchtst.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update src/transformers/models/patchtst/modeling_patchtst.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * fix tests * add add_start_docstrings * move examples to the forward's docstrings * update prepare_batch * update test * fix test_prediction_head * fix generation test * use seed to create generator * add output_hidden_states and config.num_patches * add loc and scale args in PatchTSTForPredictionOutput * edit outputs if if not return_dict * use self.share_embedding to check instead checking type. * remove seed * make style * seed is an optional int * fix test * generator device * Fix assertTrue test * swap order of items in outputs when return_dict=False. * add mask_type and random_mask_ratio to unittest * Update modeling_patchtst.py * add add_start_docstrings for regression model * make style * update model path * Edit the ValueError comment in forecast_masking * update examples * make style * fix commented code * update examples: remove config from from_pretrained call * Edit example outputs * Set default target_values to None * remove config setting in regression example * Update configuration_patchtst.py * Update configuration_patchtst.py * remove config from examples * change default d_model and ffn_dim * norm_eps default * set has_attentions to Trye and define self.seq_length = self.num_patche * update docstring * change variable mask_input to do_mask_input * fix blank space. * change logger.debug to logger.warning. * remove unused PATCHTST_INPUTS_DOCSTRING * remove all_generative_model_classes * set test_missing_keys=True * remove undefined params in the docstring. --------- Co-authored-by: nnguyen <nnguyen@us.ibm.com> Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com> Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> Co-authored-by: Nam Nguyen <namctin@gmail.com> Co-authored-by: Wesley Gifford <79663411+wgifford@users.noreply.github.com> Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2023-11-29 13:36:38 +01:00
Steven Liu	bd50402b56	[docs] Quantization (#27641 ) * first draft * benchmarks * feedback	2023-11-28 08:41:47 -08:00
Tom Aarsen	f2ad4b537b	Docs: Fix broken cross-references, i.e. `~transformer.` -> `~transformers.` (#27740 ) ~transformer. -> ~transformers.	2023-11-28 08:40:44 -08:00
Susnato Dhar	dfbd209c25	CLVP Fixes (#27547 ) * fixes * more fixes * style fix * more fix * comments	2023-11-28 17:40:01 +01:00
Yih-Dar	30e92ea323	Trigger corresponding pipeline tests if `tests/utils/tiny_model_summary.json` is modified (#27693 ) * fix --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2023-11-28 17:21:21 +01:00
Quentin Gallouédec	0b9c934575	Enforce pin memory disabling when using cpu only (#27745 ) if use_cpu: dataloader_pin_memory = False	2023-11-28 17:03:07 +01:00
Juarez Bochi	fdd86eed3b	Add madlad-400 MT models (#27471 ) * Add madlad-400 models * Add madlad-400 to the doc table * Update docs/source/en/model_doc/madlad-400.md Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Fill missing details in documentation * Update docs/source/en/model_doc/madlad-400.md Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Do not doctest madlad-400 Tests are timing out. --------- Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2023-11-28 13:19:50 +00:00
Yih-Dar	6336a7f7d6	Log a warning in `TransfoXLTokenizer.__init__` (#27721 ) * log * log --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2023-11-28 10:44:04 +01:00
Yih-Dar	93170298d1	Update tiny model creation script (#27674 ) update Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2023-11-28 10:05:34 +01:00
NielsRogge	1fb3c23b41	Add BeitBackbone (#25952 ) * First draft * Add backwards compatibility * More improvements * More improvements * Improve error message * Address comment * Add conversion script * Fix style * Update code snippet * Adddress comment * Apply suggestions from code review Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> --------- Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>	2023-11-28 08:38:32 +00:00
Yih-Dar	7a757bb694	Fix AMD Push CI not triggered (#27732 ) * fix * fix --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2023-11-28 09:30:21 +01:00
Charbel Abi Daher	2ca73e5ee3	Fixed passing scheduler-specific kwargs via TrainingArguments lr_scheduler_kwargs (#27595 ) * Fix passing scheduler-specific kwargs through TrainingArguments `lr_scheduler_kwargs` * Added test for lr_scheduler_kwargs	2023-11-28 08:33:45 +01:00
Rockerz	0864dd3beb	Translate `en/model_doc` to JP (#27264 ) * Add `model_docs` * Add * Update Model adoc * Update docs/source/ja/model_doc/bark.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/ja/model_doc/beit.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/ja/model_doc/bit.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/ja/model_doc/blenderbot.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/ja/model_doc/blenderbot-small.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * update reiew-1 * Update toctree.yml * translating docs and fixes of PR #27401 * Update docs/source/ja/model_doc/bert.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update docs/source/ja/model_doc/bert-generation.md Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com> * Update the model docs --------- Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>	2023-11-27 13:19:04 -08:00
jiaqiw09	cad1b1192b	translation main-class files to chinese (#27588 ) * translate work * update * update * update [[autodoc]] * Update callback.md --------- Co-authored-by: jiaqiw <wangjiaqi50@huawei.com>	2023-11-27 12:36:37 -08:00
Matt	74a3cebfa5	Update chat template warnings/guides (#27634 ) * Update default ChatML template * Update docs/warnings * Update docs/source/en/chat_templating.md Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Slight rework --------- Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2023-11-27 18:40:10 +00:00
Peter Pan	ce31508134	docs: replace torch.distributed.run by torchrun (#27528 ) * docs: replace torch.distributed.run by torchrun `transformers` now officially support pytorch >= 1.10. The entrypoint `torchrun`` is present from 1.10 onwards. Signed-off-by: Peter Pan <Peter.Pan@daocloud.io> * Update src/transformers/trainer.py with @ArthurZucker's suggestion Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> --------- Signed-off-by: Peter Pan <Peter.Pan@daocloud.io> Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>	2023-11-27 16:26:33 +00:00
NielsRogge	c832bcb812	Fix owlv2 code snippet (#27698 ) * Fix code snippet * Improve code snippet	2023-11-27 16:29:07 +01:00
Yixiao Yuan	334a6d18a1	Modify group_sub_entities in TokenClassification Pipeline to support label with "-" (#27325 ) * fix group_sub_entities bug * add space	2023-11-27 15:25:46 +00:00
NielsRogge	59499bbe8b	Update forward signature test for vision models (#27681 ) * Update forward signature * Empty-Commit	2023-11-27 15:48:17 +01:00
jiqing-feng	1d7f406e19	fix assisted decoding assistant model inputs (#27503 ) * fix assisted decoding attention_cat * fix attention_mask for assisted decoding * fix attention_mask len * fix attn len * Use a more clean way to prepare assistant models inputs * fix param meaning * fix param name * fix assistant model inputs * update token type ids * fix assistant kwargs copy * add encoder-decoder tests of assisted decoding * check if assistant kwargs contains updated keys * revert test * fix whisper tests * fix assistant kwargs * revert whisper test * delete _extend funcs	2023-11-27 14:23:54 +00:00
yhshin11	307cf3a2ab	Fix oneformer instance segmentation RuntimeError (#27725 )	2023-11-27 14:59:59 +01:00
Yanan Xie	b09912c8f4	Fix mistral generate for long prompt / response (#27548 ) * Fix mistral generate for long prompt / response * Add unit test * fix linter * fix linter * fix test * add assisted generation test for mistral and load the model in 4 bit + fa2	2023-11-27 10:18:41 +01:00
Lysandre Debut	27b752bcf1	Reorder the code on the Hub to explicit that sharing on the Hub isn't a requirement (#27691 ) Reorder	2023-11-27 09:38:18 +01:00
Arthur	5c30dd40e7	fix warning (#27689 )	2023-11-27 09:14:40 +01:00
Yih-Dar	e11e26df93	Fix Past CI (#27696 ) fix Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2023-11-27 09:11:58 +01:00
Ilya Gusev	f70db28322	Fix sliding_window hasattr in Mistral (#27041 ) * Fix sliding_window hasattr in Mistral * hasattr -> getattr for sliding_window in Mistral --------- Co-authored-by: Ilya Gusev <ilya.gusev@booking.com>	2023-11-26 16:28:37 +01:00
Yih-Dar	35551f9a0f	Fix `TVPModelTest` (#27695 ) * fix * fix * fix * fix * fix --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>	2023-11-24 19:47:50 +01:00
Chi	29c94808ea	Successfully Resolved The ZeroDivisionError Exception. (#27524 ) * Successfully resolved the ZeroDivisionError exception in the utils.notebook.y file. * Now I update little code mentioned by Peter * Using Black package to reformat my file * Now I using ruff libary to reformated my file	2023-11-24 16:55:08 +00:00
fxmarty	c13a43aaf2	Reflect RoCm support in the documentation (#27636 ) * reflect RoCm support in the documentation * Update docs/source/en/main_classes/trainer.md Co-authored-by: Lysandre Debut <hi@lysand.re> * fix review comments * use ROCm instead of RoCm --------- Co-authored-by: Lysandre Debut <hi@lysand.re>	2023-11-25 00:59:17 +09:00
Arthur	a6d178e238	[`DocString`] Support a revision in the docstring `add_code_sample_docstrings` to facilitate integrations (#27645 ) * initial commit * dummy changes * style * Update src/transformers/utils/doc.py Co-authored-by: Alex McKinney <44398246+vvvm23@users.noreply.github.com> * nits * nit use ` if re.match(r'^refs/pr/\d', revision):` restrict * nit * test the doc vuilder * wow * oke the order was wrong --------- Co-authored-by: Alex McKinney <44398246+vvvm23@users.noreply.github.com>	2023-11-24 16:30:05 +01:00
Anirudh Haritas Murali	2098d343cc	Fix semantic error in evaluation section (#27675 ) Change "convert predictions to logits" to "convert logits to predictions" to fix semantic error in the evaluation section. Logits need to be converted to predictions to evaluate the accuracy, not the other way round	2023-11-24 12:41:16 +01:00

1 2 3 4 5 ...

14621 Commits