Commits · 85fde09c97213bf7e8625f83096bb2a9e183f987 · zhusg / transformers-new

16 Nov, 2023 3 commits

[`pytest`] Avoid flash attn test marker warning (#27509) · 85fde09c
Arthur authored 1 year ago
```
add flash attn markers
```
85fde09c
Support ONNX export for causal LM sequence classifiers (#27450) · 1394e08c
Dean Wyatte authored 1 year ago
```
support onnx for causal lm sequence classification
```
1394e08c

translate model.md to chinese (#27518) · 06343b06

Hz, Ji authored 1 year ago


* translate model.md to chinese

* apply review suggestion

Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

---------

Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

06343b06

15 Nov, 2023 15 commits

Fix offload disk for loading derivated model checkpoint into base model (#27253) · 1ac599d9
Marc Sun authored 1 year ago
```
* fix

* style

* add test
```
1ac599d9

Fix bug for T5x to PyTorch convert script with varying encoder and decoder layers (#27448) · b71c38a0

JiangZhongqing authored 1 year ago

* Fix bug in handling varying encoder and decoder layers

This commit resolves an issue where the script failed to convert T5x models to PyTorch models when the number of decoder layers differed from the number of encoder layers.  I've addressed this issue by passing an additional 'num_decoder_layers' parameter to the relevant function.

* Fix bug in handling varying encoder and decoder layers

b71c38a0

Incorrect setting for num_beams in translation and summarization examples (#27519) · 2e72bbab

Matt authored 1 year ago


* Remove the torch main_process_first context manager from TF examples

* Correctly set num_beams=1 in our examples, and add a guard in GenerationConfig.validate()

* Update src/transformers/generation/configuration_utils.py

Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

---------

Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

2e72bbab

Fixing the failure of models without max_position_embeddings attribute. (#27499) · e6522e49

Adam Louly authored 1 year ago

fix max pos issue

Co-authored-by: Adam Louly <adamlouly@microsoft.com@orttrainingdev9.d32nl1ml4oruzj4qz3bqlggovf.px.internal.cloudapp.net>

e6522e49

Translating `en/model_doc` docs to Japanese. (#27401) · a0633c44

Yuki-Imajuku authored 1 year ago


* update _toctree.yml & add albert-autoformer

* Fixed typo in docs/source/ja/model_doc/audio-spectrogram-transformer.md

Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Delete duplicated sentence docs/source/ja/model_doc/autoformer.md

Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

* Reflect reviews

* delete untranslated models from toctree

* delete all comments

* add abstract translation

---------

Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>

a0633c44

Fix wav2vec2 params (#27515) · a85ea4b1
Zach Mueller authored 1 year ago
```
Fix test
```
a85ea4b1

[ `PretrainedConfig`] Improve messaging (#27438) · 48ba1e07

Arthur authored 1 year ago


* import hf error

* nits

* fixup

* catch the error at the correct place

* style

* improve message a tiny bit

* Update src/transformers/utils/hub.py

Co-authored-by: Lucain <lucainp@gmail.com>

* add a test

---------

Co-authored-by: Lucain <lucainp@gmail.com>

48ba1e07

🚨

Fix beam score calculation issue for decoder-only models (#27351) · 453079c7

Xin Qiu authored 1 year ago


* Fix beam score calculation issue for decoder-only models

* Update beam search test and fix code quality issue

* Fix beam_sample, group_beam_search and constrained_beam_search

* Split test for pytorch and TF, add documentation

---------

Co-authored-by: Xin Qiu <xin.qiu@sentient.ai>

453079c7

[`tokenizers`] update `tokenizers` version pin (#27494) · 3d1a7bf4

Arthur authored 1 year ago


* update `tokenizers` version pin

* force tokenizers>=0.15

* use  0.14

Co-authored-by: Lysandre <lysandre@huggingface.co>

---------

Co-authored-by: Lysandre <lysandre@huggingface.co>

3d1a7bf4

Make some jobs run on the GitHub Actions runners (#27512) · 64e21ca2
Yih-Dar authored 1 year ago
```
fix

Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
```
64e21ca2

[`CircleCI`] skip test_assisted_decoding_sample for everyone (#27511) · 1e0e2dd3

Arthur authored 1 year ago

* skip 4 tests

* nits

* style

* wow it's not my day

* skip new failing tests

* style

* skip for NLLB MoE as well

* skip `test_assisted_decoding_sample` for everyone

1e0e2dd3

Update spelling mistake (#27506) · 7ddb21b4
Phyzer authored 1 year ago
```
thoroughly was misspelled thouroughly
```
7ddb21b4
[Table Transformer] Add Transformers-native checkpoints (#26928) · 72f531ab
NielsRogge authored 1 year ago
```
* Improve conversion scripts

* Fix paths

* Fix style
```
72f531ab

[Fuyu] Add tests (#27001) · cc0dc24b

NielsRogge authored 1 year ago

* Add tests

* Add integration test

* More improvements

* Fix tests

* Fix style

* Skip gradient checkpointing tests

* Update script

* Remove scripts

* Remove Fuyu from auto mapping

* Fix integration test

* More improvements

* Remove file

* Add Fuyu to slow documentation tests

* Address comments

* Clarify comment

cc0dc24b

[`CI-test_torch`] skip test_tf_from_pt_safetensors and `test_assisted_decoding_sample` (#27508) · 186c0775
Arthur authored 1 year ago
```
* skip 4 tests

* nits

* style

* wow it's not my day

* skip new failing tests

* style

* skip for NLLB MoE as well
```
186c0775

14 Nov, 2023 16 commits

Track the number of tokens seen to metrics (#27274) · 2fc33ebe

Zach Mueller authored 1 year ago


* Add tokens seen

* Address comments, add to TrainingArgs

* Update log

* Apply suggestions from code review

Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* Use self.args

* Fix docstring

Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

---------

Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

2fc33ebe

Update processor mapping for hub snippets (#27477) · 303c1d69
amyeroberts authored 1 year ago

303c1d69

Have seq2seq just use gather (#27025) · 067c4a31

Zach Mueller authored 1 year ago


* Have seq2seq just use gather

* Change

* Reset after

* Make slow

* Apply suggestions from code review

Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

* Clean

* Simplify and just use gather

* Update tests/trainer/test_trainer_seq2seq.py

Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* gather always for seq2seq

---------

Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>
Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

067c4a31

Minor type annotation fix (#27276) · 250032e9
Costa Huang authored 1 year ago
```
* Minor type annotation fix

* Trigger Build
```
250032e9
Generate: `GenerationConfig.from_pretrained` can return unused kwargs (#27488) · a53a0c51
Joao Gante authored 1 year ago

a53a0c51

Update and reorder docs for chat templates (#27443) · 5468ab35

Matt authored 1 year ago

* Update and reorder docs for chat templates

* Fix Mistral docstring

* Add section link and small fixes

* Remove unneeded line in Mistral example

* Add comment on saving memory

* Fix generation prompts linl

* Fix code block languages

5468ab35

Generate: fix `ExponentialDecayLengthPenalty` doctest (#27485) · fe472b1d
Joao Gante authored 1 year ago
```
fix exponential doctest
```
fe472b1d
translate hpo_train.md and perf_hardware.md to chinese (#27431) · 73bc0c9e
jiaqiw09 authored 1 year ago
```
* translate

* translate

* update
```
73bc0c9e
Revert "[time series] Add PatchTST (#25927)" (#27486) · 78f6ed6c
amyeroberts authored 1 year ago
```
The model was merged before final review and approval.

This reverts commit 2ac5b932.
```
78f6ed6c
[Whisper] Fix pipeline test (#27442) · a4616c67
Sanchit Gandhi authored 1 year ago

a4616c67

Clap processor: remove wasteful np.stack operations (#27454) · b86c54d9

Max Bain authored 1 year ago

remove wasteful np.stack

Np.stack on large 1-D tensor, causing ~0.5s processing time on short audio (<10s). Compared to 0.02s for medium length audio

b86c54d9

Add speecht5 batch generation and fix wrong attention mask when padding (#25943) · 4309abed

Sihan Chen authored 1 year ago

* fix speecht5 wrong attention mask when padding

* enable batch generation and add parameter attention_mask

* fix doc

* fix format

* batch postnet inputs, return batched lengths, and consistent to old api

* fix format

* fix format

* fix the format

* fix doc-builder error

* add test, cross attention and docstring

* optimize code based on reviews

* docbuild

* refine

* not skip slow test

* add consistent dropout for batching

* loose atol

* add another test regarding to the consistency of vocoder

* fix format

* refactor

* add return_concrete_lengths as parameter for consistency w/wo batching

* fix review issues

* fix cross_attention issue

4309abed

Fix M4T weights tying (#27395) · ee4fb326
Yoach Lacombe authored 1 year ago
```
fix seamless m4t weights tying
```
ee4fb326
[`CI-test_torch`] skip `test_tf_from_pt_safetensors` for 4 models (#27481) · e107ae36
Arthur authored 1 year ago
```
* skip 4 tests

* nits

* style

* wow it's not my day
```
e107ae36

[`Peft`] `modules_to_save` support for peft integration (#27466) · d71fa9f6

Younes Belkada authored 1 year ago


* `modules_to_save` support for peft integration

* Update docs/source/en/peft.md

Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* slightly elaborate test

---------

Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

d71fa9f6

Fix FA2 import + deprecation cycle (#27330) · 721d1c8c
Marc Sun authored 1 year ago
```
* put back import

* switch to logger.warnings instead
```
721d1c8c

13 Nov, 2023 6 commits

[time series] Add PatchTST (#25927) · 2ac5b932

Gift Sinthong authored 1 year ago


* Initial commit of PatchTST model classes

Co-authored-by: Phanwadee Sinthong <phsinthong@gmail.com>
Co-authored-by: Nam Nguyen <namctin@gmail.com>
Co-authored-by: Vijay Ekambaram <vijaykr.e@gmail.com>
Co-authored-by: Ngoc Diep Do <55230119+diepi@users.noreply.github.com>
Co-authored-by: Wesley Gifford <79663411+wgifford@users.noreply.github.com>

* Add PatchTSTForPretraining

* update to include classification

Co-authored-by: Phanwadee Sinthong <phsinthong@gmail.com>
Co-authored-by: Nam Nguyen <namctin@gmail.com>
Co-authored-by: Vijay Ekambaram <vijaykr.e@gmail.com>
Co-authored-by: Ngoc Diep Do <55230119+diepi@users.noreply.github.com>
Co-authored-by: Wesley Gifford <79663411+wgifford@users.noreply.github.com>

* clean up auto files

* Add PatchTSTForPrediction

* Fix relative import

* Replace original PatchTSTEncoder with ChannelAttentionPatchTSTEncoder

* temporary adding absolute path + add PatchTSTForForecasting class

* Update base PatchTSTModel + Unittest

* Update ForecastHead to use the config class

* edit cv_random_masking, add mask to model output

* Update configuration_patchtst.py

* add masked_loss to the pretraining

* add PatchEmbeddings

* Update configuration_patchtst.py

* edit loss which considers mask in the pretraining

* remove patch_last option

* Add commits from internal repo

* Update ForecastHead

* Add model weight initilization + unittest

* Update PatchTST unittest to use local import

* PatchTST integration tests for pretraining and prediction

* Added PatchTSTForRegression + update unittest to include label generation

* Revert unrelated model test file

* Combine similar output classes

* update PredictionHead

* Update configuration_patchtst.py

* Add Revin

* small edit to PatchTSTModelOutputWithNoAttention

* Update modeling_patchtst.py

* Updating integration test for forecasting

* Fix unittest after class structure changed

* docstring updates

* change input_size to num_input_channels

* more formatting

* Remove some unused params

* Add a comment for pretrained models

* add channel_attention option

add channel_attention option and remove unused positional encoders.

* Update PatchTST models to use HF's MultiHeadAttention module

* Update paper + github urls

* Fix hidden_state return value

* Update integration test to use PatchTSTForForecasting

* Adding dataclass decorator for model output classes

* Run fixup script

* Rename model repos for integration test

* edit argument explanation

* change individual option to shared_projection

* style

* Rename integration test + import cleanup

* Fix outpu_hidden_states return value

* removed unused mode

* added std, mean and nops scaler

* add initial distributional loss for predition

* fix typo in docs

* add generate function

* formatting

* add num_parallel_samples

* Fix a typo

* copy weighted_average function, edit PredictionHead

* edit PredictionHead

* add distribution head to forecasting

* formatting

* Add generate function for forecasting

* Add generate function to prediction task

* formatting

* use argsort

* add past_observed_mask ordering

* fix arguments

* docs

* add back test_model_outputs_equivalence test

* formatting

* cleanup

* formatting

* use ACT2CLS

* formatting

* fix add_start_docstrings decorator

* add distribution head and generate function to regression task

add distribution head and generate function to regression task. Also made add PatchTSTForForecastingOutput,  PatchTSTForRegressionOutput.

* add distribution head and generate function to regression task

add distribution head and generate function to regression task. Also made add PatchTSTForForecastingOutput,  PatchTSTForRegressionOutput.

* fix typos

* add forecast_masking

* fixed tests

* use set_seed

* fix doc test

* formatting

* Update docs/source/en/model_doc/patchtst.md

Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>

* better var names

* rename PatchTSTTranspose

* fix argument names and docs string

* remove compute_num_patches and unused class

* remove assert

* renamed to PatchTSTMasking

* use num_labels for classification

* use num_labels

* use default num_labels from super class

* move model_type after docstring

* renamed PatchTSTForMaskPretraining

* bs -> batch_size

* more review fixes

* use hidden_state

* rename encoder layer and block class

* remove commented seed_number

* edit docstring

* Add docstring

* formatting

* use past_observed_mask

* doc suggestion

* make fix-copies

* use Args:

* add docstring

* add docstring

* change some variable names and add PatchTST before some class names

* formatting

* fix argument types

* fix tests

* change x variable to patch_input

* format

* formatting

* fix-copies

* Update tests/models/patchtst/test_modeling_patchtst.py

Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>

* move loss to forward

* Update src/transformers/models/patchtst/modeling_patchtst.py

Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>

* Update src/transformers/models/patchtst/modeling_patchtst.py

Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>

* Update src/transformers/models/patchtst/modeling_patchtst.py

Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>

* Update src/transformers/models/patchtst/modeling_patchtst.py

Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>

* Update src/transformers/models/patchtst/modeling_patchtst.py

Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>

* formatting

* fix a bug when pre_norm is set to True

* output_hidden_states is set to False as default

* set pre_norm=True as default

* format docstring

* format

* output_hidden_states is None by default

* add missing docs

* better var names

* docstring: remove default to False in output_hidden_states

* change labels name to target_values in regression task

* format

* fix tests

* change to forecast_mask_ratios and random_mask_ratio

* change mask names

* change future_values to target_values param in the prediction class

* remove nn.Sequential and make PatchTSTBatchNorm class

* black

* fix argument name for prediction

* add output_attentions option

* add output_attentions to PatchTSTEncoder

* formatting

* Add attention output option to all classes

* Remove PatchTSTEncoderBlock

* create PatchTSTEmbedding class

* use config in PatchTSTPatchify

* Use config in PatchTSTMasking class

* add channel_attn_weights

* Add PatchTSTScaler class

* add output_attentions arg to test function

* format

* Update doc with image patchtst.md

* fix-copies

* rename Forecast <-> Prediction

* change name of a few parameters to match with PatchTSMixer.

* Remove *ForForecasting class to match with other time series models.

* make style

* Remove PatchTSTForForecasting in the test

* remove PatchTSTForForecastingOutput class

* change test_forecast_head to test_prediction_head

* style

* fix docs

* fix tests

* change num_labels to num_targets

* Remove PatchTSTTranspose

* remove arguments in PatchTSTMeanScaler

* remove arguments in PatchTSTStdScaler

* add config as an argument to all the scaler classes

* reformat

* Add norm_eps for batchnorm and layernorm

* reformat.

* reformat

* edit docstring

* update docstring

* change variable name pooling to pooling_type

* fix output_hidden_states as tuple

* fix bug when calling PatchTSTBatchNorm

* change stride to patch_stride

* create PatchTSTPositionalEncoding class and restructure the PatchTSTEncoder

* formatting

* initialize scalers with configs

* edit output_hidden_states

* style

* fix forecast_mask_patches doc string

---------

Co-authored-by: Gift Sinthong <gift.sinthong@ibm.com>
Co-authored-by: Nam Nguyen <namctin@gmail.com>
Co-authored-by: Vijay Ekambaram <vijaykr.e@gmail.com>
Co-authored-by: Ngoc Diep Do <55230119+diepi@users.noreply.github.com>
Co-authored-by: Wesley Gifford <79663411+wgifford@users.noreply.github.com>
Co-authored-by: Wesley M. Gifford <wmgifford@us.ibm.com>
Co-authored-by: nnguyen <nnguyen@us.ibm.com>
Co-authored-by: Ngoc Diep Do <diiepy@gmail.com>
Co-authored-by: Kashif Rasul <kashif.rasul@gmail.com>
Co-authored-by: NielsRogge <48327001+NielsRogge@users.noreply.github.com>
Co-authored-by: Patrick von Platen <patrick.v.platen@gmail.com>

2ac5b932

Fixed typo in pipelines.md documentation (#27455) · 8017a590
adismort14 authored 1 year ago
```
Update pipelines.md
```
8017a590
Perf torch compile (#27422) · eb79b55b
jiaqiw09 authored 1 year ago
```
* translate perrf_torch_compile.md

* translate tf_xla.md

* update
```
eb79b55b
[`AWQ` ] Addresses TODO for awq tests (#27467) · 7b139023
Younes Belkada authored 1 year ago
```
addresses todo for awq tests
```
7b139023

Fix Falcon tokenizer loading in pipeline (#27316) · 04af4b90

Matt authored 1 year ago

* Improve pipeline tokenizer loading and hope nothing breaks

* Let's try a hacky solution

* Revert the changes to init

* Add a falcon hack to the automapping

* Add a falcon hack to the automapping

04af4b90

Add version check for Jinja (#27403) · 1af766e1

Matt authored 1 year ago


* Add version check for Jinja

* Update src/transformers/tokenization_utils_base.py

Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

* make fixup

---------

Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com>

1af766e1