MoE (Mixture of Experts) layer

Appears in 1 paper · 1 tutorial

A drop-in replacement for the FFN sub-layer in a Transformer.

As used in Paper 09 — Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer →

A drop-in replacement for the FFN sub-layer in a Transformer. Contains n expert networks and a gating network. For each token, routes to top-k experts and outputs a weighted sum of their outputs. Keeps attention sub-layers dense and shared across all tokens.

As used in Fine-Tuning & Model Customization →

An architecture with many expert sub-networks and a router that activates a few per token. (M15)