BERT
BERT introduced bidirectional pre-training.
Why it matters
It established a reusable bidirectional pre-training recipe for language tasks.
Its reported benchmarks depend on the model and evaluation tasks described in the paper.
Research · October 11, 2018
BERT introduced bidirectional pre-training.
It established a reusable bidirectional pre-training recipe for language tasks.
Its reported benchmarks depend on the model and evaluation tasks described in the paper.
arXiv · Published October 11, 2018
BERT uses a bidirectional Transformer encoder, explicitly described in the paper.
BERTA different bet on the same mechanism: generation instead of bidirectional encoding.
Improving Language Understanding by Generative Pre-TrainingMasked self-supervision is taken from the approach shown on text.
The aggregate score over nine tasks was the number pretrained models were measured by half a year later. No source in this record asserts causation.
The model takes the same transformer settings as BERT and is measured against it task by task.
ERNIE 2.0: A Continual Pre-training Framework for Language Understanding (arXiv:1907.12412v1)