What is the difference between an encoder and a decoder in transformers?
PICTURE THIS: A SENTENCE BECOMES TOKENS
The model does not read letters like humans. It reads these pieces, then predicts the next one.
Simple meaning
An encoder builds bidirectional representations of the full input.
WHY — Transformers instead of guessing?
Why interviewers care about Transformers:
contrast on Transformers, not two memorised paragraphs.
the developer, then one case where picking wrong hurts.
Name the idea, why it exists, then one short example.
End with when you use it and one common pitfall.
STEPS — What happens with tokens?
Before the model can read a sentence, it goes through these steps:
- 1An encoder builds bidirectional
representations of the full input.
- 2A decoder generates tokens
left to right and can only see past tokens plus optional encoder output.
- 3Embeddings
BERT is encoder-style
- 4Context mix
GPT is decoder-style
- 5Next token
T5 uses both.
- 6Decode
IDs turn back into readable text.
EXAMPLE — See it in action
Let's see how a real sentence is tokenized (tokens may vary by model):
Note: Actual tokens and IDs depend on the tokenizer (e.g., GPT, Llama, etc.).
Key takeaway
An encoder builds bidirectional representations of the full input. A decoder generates tokens left to right and can only see past tokens plus optional encoder output.