transformers

mirror of https://github.com/huggingface/transformers.git synced 2025-07-07 23:00:08 +06:00

History

Nicolas Patry db9dd09cf9 Adding `AutomaticSpeechRecognitionPipeline`. (#11337 ) * Adding `AutomaticSpeechRecognitionPipeline`. - Because we added everything to enable this pipeline, we probably should add it to `transformers`. - This PR tries to limit the scope and focuses only on the pipeline part (what should go in, and out). - The tests are very specific for S2T and Wav2vec2 to make sure both architectures are supported by the pipeline. We don't use the mixin for tests right now, because that requires more work in the `pipeline` function (will be done in a follow up PR). - Unsure about the "helper" function `ffmpeg_read`. It makes a lot of sense from a user perspective, it does not add any additional dependencies (as in hard dependency, because users can always use their own load mechanism). Meanwhile, it feels slightly clunky to have so much optional preprocessing. - The pipeline is not done to support streaming audio right now. Future work: - Add `automatic-speech-recognition` as a `task`. And add the FeatureExtractor.from_pretrained within `pipeline` function. - Add small models within tests - Add the Mixin to tests. - Make the logic between ForCTC vs ForConditionalGeneration better. * Update tests/test_pipelines_automatic_speech_recognition.py Co-authored-by: Lysandre Debut <lysandre@huggingface.co> * Adding docs + main import + type checking + LICENSE. * Doc style !. * Fixing TYPE_HINT. * Specifying waveform shape in the docs. * Adding asserts + specify in the documentation the shape of the input np.ndarray. * Update src/transformers/pipelines/automatic_speech_recognition.py Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com> * Adding require to tests + move the `feature_extractor` doc. Co-authored-by: Lysandre Debut <lysandre@huggingface.co> Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>		2021-04-30 11:54:08 +02:00
..
callback.rst	Add example for registering callbacks with trainers (#10928 )	2021-04-05 12:27:23 -04:00
configuration.rst	Copyright (#8970 )	2020-12-07 18:36:34 -05:00
data_collator.rst	Doc check: a bit of clean up (#11224 )	2021-04-13 12:14:25 -04:00
feature_extractor.rst	Add ImageFeatureExtractionMixin (#10905 )	2021-03-26 11:23:56 -04:00
logging.rst	Logging propagation (#10092 )	2021-02-09 10:27:49 -05:00
model.rst	Trainer push to hub (#11328 )	2021-04-23 09:17:37 -04:00
optimizer_schedules.rst	Seq2seq trainer (#9241 )	2020-12-22 11:33:44 -05:00
output.rst	update QuickTour docs to reflect model output object (#11462 )	2021-04-26 22:18:37 -04:00
pipelines.rst	Adding `AutomaticSpeechRecognitionPipeline`. (#11337 )	2021-04-30 11:54:08 +02:00
processors.rst	Examples reorg (#11350 )	2021-04-21 11:11:20 -04:00
tokenizer.rst	Documentation about loading a fast tokenizer within Transformers (#11029 )	2021-04-05 10:51:16 -04:00
trainer.rst	[Deepspeed] ZeRO-Infinity integration plus config revamp (#11418 )	2021-04-26 10:40:32 -07:00