mirror of https://github.com/huggingface/transformers.git synced 2025-07-04 05:10:06 +06:00

History

Albert Villanova del Moral a14b055b65 Pass datasets trust_remote_code (#31406 ) * Pass datasets trust_remote_code * Pass trust_remote_code in more tests * Add trust_remote_dataset_code arg to some tests * Revert "Temporarily pin datasets upper version to fix CI" This reverts commit `b7672826ca`. * Pass trust_remote_code in librispeech_asr_dummy docstrings * Revert "Pin datasets<2.20.0 for examples" This reverts commit `833fc17a3e`. * Pass trust_remote_code to all examples * Revert "Add trust_remote_dataset_code arg to some tests" to research_projects * Pass trust_remote_code to tests * Pass trust_remote_code to docstrings * Fix flax examples tests requirements * Pass trust_remote_dataset_code arg to tests * Replace trust_remote_dataset_code with trust_remote_code in one example * Fix duplicate trust_remote_code * Replace args.trust_remote_dataset_code with args.trust_remote_code * Replace trust_remote_dataset_code with trust_remote_code in parser * Replace trust_remote_dataset_code with trust_remote_code in dataclasses * Replace trust_remote_dataset_code with trust_remote_code arg		2024-06-17 17:29:13 +01:00
..
README.md	Update all references to canonical models (#29001 )	2024-02-16 08:16:58 +01:00
requirements.txt	Update reqs to include min gather_for_metrics Accelerate version (#20242 )	2022-11-15 13:28:00 -05:00
run_no_trainer.sh	Examples reorg (#11350 )	2021-04-21 11:11:20 -04:00
run_swag_no_trainer.py	Pass datasets trust_remote_code (#31406 )	2024-06-17 17:29:13 +01:00
run_swag.py	v4.42.dev.0	2024-05-17 17:30:41 +02:00

README.md

Multiple Choice

Fine-tuning on SWAG with the Trainer

run_swag allows you to fine-tune any model from our hub (as long as its architecture as a ForMultipleChoice version in the library) on the SWAG dataset or your own csv/jsonlines files as long as they are structured the same way. To make it works on another dataset, you will need to tweak the preprocess_function inside the script.

python examples/multiple-choice/run_swag.py \
--model_name_or_path FacebookAI/roberta-base \
--do_train \
--do_eval \
--learning_rate 5e-5 \
--num_train_epochs 3 \
--output_dir /tmp/swag_base \
--per_device_eval_batch_size=16 \
--per_device_train_batch_size=16 \
--overwrite_output

Training with the defined hyper-parameters yields the following results:

***** Eval results *****
eval_acc = 0.8338998300509847
eval_loss = 0.44457291918821606

With Accelerate

Based on the script run_swag_no_trainer.py.

Like run_swag.py, this script allows you to fine-tune any of the models on the hub (as long as its architecture as a ForMultipleChoice version in the library) on the SWAG dataset or your own data in a csv or a JSON file. The main difference is that this script exposes the bare training loop, to allow you to quickly experiment and add any customization you would like.

It offers less options than the script with Trainer (but you can easily change the options for the optimizer or the dataloaders directly in the script) but still run in a distributed setup, on TPU and supports mixed precision by the mean of the 🤗 Accelerate library. You can use the script normally after installing it:

pip install git+https://github.com/huggingface/accelerate

then

export DATASET_NAME=swag

python run_swag_no_trainer.py \
  --model_name_or_path google-bert/bert-base-cased \
  --dataset_name $DATASET_NAME \
  --max_seq_length 128 \
  --per_device_train_batch_size 32 \
  --learning_rate 2e-5 \
  --num_train_epochs 3 \
  --output_dir /tmp/$DATASET_NAME/

You can then use your usual launchers to run in it in a distributed environment, but the easiest way is to run

accelerate config

and reply to the questions asked. Then

accelerate test

that will check everything is ready for training. Finally, you can launch training with

export DATASET_NAME=swag

accelerate launch run_swag_no_trainer.py \
  --model_name_or_path google-bert/bert-base-cased \
  --dataset_name $DATASET_NAME \
  --max_seq_length 128 \
  --per_device_train_batch_size 32 \
  --learning_rate 2e-5 \
  --num_train_epochs 3 \
  --output_dir /tmp/$DATASET_NAME/

This command is the same and will work for:

a CPU-only setup
a setup with one GPU
a distributed training with several GPUs (single or multi node)
a training on TPUs

Note that this library is in alpha release so your feedback is more than welcome if you encounter any problem using it.