Ainiketan ainiketan.in
Papers Applied AI Dictionary This Week Learning Paths ∑ Playground
Papers Applied AI Dictionary This Week Learning Paths ∑ Playground
← Dictionary / Transformer Decoder

Transformer Decoder

Appears in 1 paper

The architecture used in GPT models: a stack of self-attention and feedforward layers that process tokens left-to-right (causally).

As used in Paper 13 — Scaling Laws for Neural Language Models →

The architecture used in GPT models: a stack of self-attention and feedforward layers that process tokens left-to-right (causally). Each token can only attend to previous tokens.

Appears in

Paper 13 — Scaling Laws for Neural Language Models →
Browse Dictionary
← All terms A–Z
Share
WhatsApp
Ainiketan

Where India learns AI — deeply, freely, together.

जहाँ हर जिज्ञासु AI सीखे — खुलकर, गहराई से, साथ में।

Free forever No ads No login Open source

Learn

All 24 Papers Math Tutorials Dictionary Learning Paths This Week in AI

Community

Student Journal Soon Paper Club Soon Research Questions Soon Mentor Network Soon Teacher Packs Soon

Site

About Scholarship Fund Impact Corrections Terms & Copyright
Weekly digest

5 things in AI every week. Plain English. Free.

© 2026 Ainiketan · Built for India, for free, forever · Suggest a correction

Content license: CC BY 4.0 · Hosted on Vercel · Privacy-friendly analytics (no cookies)

All summaries are original writing by Ainiketan — we link to sources and do not reproduce copyrighted text. Copyright concerns: askainiketan@gmail.com · Terms & Copyright