A vector quantizer can expose thousands of codebook entries while repeatedly selecting only a small subset. The configured vocabulary then overstates the discrete capacity that is actually used. This state is often called codebook collapse: some entries receive frequent assignments, while others become inactive or nearly inactive.
The failure is not a shortage of parameters by itself. It comes from the interaction between encoder outputs, the assignment rule, and the procedure used to update code vectors. Once assignments concentrate around a subset of entries, unused entries may receive little or no signal that would move them toward regions occupied by encoder outputs.
Nearest-code assignment creates a discrete routing boundary
Consider an encoder output z and a codebook containing vectors e_1 ... e_K. A common hard quantizer selects the nearest entry:
k* = argmin_k ||z - e_k||²
q(z) = e_k*This operation partitions representation space into regions associated with codebook entries. An entry is active for a batch only if at least one encoder output falls into its assignment region.
The configured value K therefore does not guarantee K useful symbols. If encoder outputs occupy a narrow part of the space, or if several code vectors are poorly positioned, assignments can concentrate on fewer entries. A large codebook can still have low effective utilization.
Hard assignment also creates a feedback effect. An entry that receives assignments can be updated toward observed encoder outputs. An entry that receives none may lack a direct assignment-driven update. The exact update path depends on the quantizer implementation: gradient-based codebook updates and exponential-moving-average updates do not have identical dynamics.
Utilization is separate from reconstruction quality
A reconstruction objective can improve even when code usage is uneven. The decoder may represent the observed data adequately using a limited subset of codes, especially when the active codes and decoder have enough capacity for the training distribution.
For that reason, reconstruction loss alone does not reveal whether the discrete bottleneck is using the intended vocabulary. Code usage needs a separate measurement. For a dataset or evaluation stream, assignment counts can be written as:
n_k = number of assignments to code k
p_k = n_k / sum_j n_jThe number of codes with n_k > 0 gives a direct occupancy count over the measurement window. Entropy provides another view:
H(p) = -sum_k p_k log p_kA higher entropy indicates a more spread assignment distribution under the chosen sample and measurement window, but it is not a quality metric by itself. Uniform use can be undesirable when the underlying data distribution is strongly non-uniform. The useful question is whether the observed occupancy matches the role expected from the discrete representation.
Inactive entries can persist under assignment-driven updates
Suppose a code vector sits far from every encoder output. Under nearest-neighbor assignment, it receives no samples. If its update is computed only from assigned samples, the entry has no local data signal pulling it toward an occupied region.
This creates a practical asymmetry. Active entries keep adapting because they keep receiving assignments. Inactive entries can remain outside the encoder distribution unless another mechanism relocates or reinitializes them.
EMA-based quantizers make this boundary explicit. Their code vectors are typically updated from running assignment counts and sums. An entry with negligible count has weak evidence for a useful centroid. Implementations can handle that state differently, so behavior around empty entries should be treated as an implementation choice rather than a universal property of vector quantization.
Gradient-based formulations have a related issue. The nearest-code decision is discrete, and training commonly uses an estimator or separated loss terms so gradients can update the encoder and codebook. The resulting dynamics depend on those loss definitions and update rules. The existence of a code vector does not ensure that optimization will route data through it.
Encoder drift can move the assignment problem
Codebook state and encoder state evolve together during training. Even a well-spread initial codebook can become poorly matched if the encoder distribution moves. Conversely, code vectors that track early encoder outputs can become concentrated in regions that later receive less mass.
This coupling means utilization should be interpreted over time, not from one final snapshot alone. A code that is inactive in one interval may become active later, and a code that was initially common may disappear from assignments. Persistent inactivity is more informative than a single empty batch.
Batch composition also matters. Rare semantic or acoustic patterns may legitimately activate only a few entries and may not appear in every batch. Declaring an entry dead from a short window can confuse sparse but valid use with persistent collapse.
Reinitialization changes the optimization state
One response to persistent inactivity is to replace unused code vectors with values derived from current encoder outputs. That can put the entries back into occupied representation regions, giving them a chance to receive future assignments.
This is not a neutral bookkeeping operation. Reinitialization changes the quantizer state and can change assignments immediately. The replacement rule, inactivity threshold, sampling procedure, and update timing therefore become part of the optimization design.
Other designs alter the assignment pressure itself, adjust codebook or commitment terms, change initialization, or use quantizers with different update mechanics. None of these choices guarantees broad utilization across every data distribution. A method that increases occupancy can also change reconstruction behavior or the semantics carried by individual codes.
The narrow implementation boundary is useful: codebook size specifies available discrete entries, while assignment statistics reveal the entries actually used. Treating those two quantities as interchangeable hides collapse precisely where the configured capacity appears healthy.