FRAMES STORE QUERY PAST KEYS/VALS SEPARATED AUDIO FRAMES CAUSAL CONV CACHE MEMORY CAUSAL ATTN SPEECH STREAMS Incoming raw audio stream divided into overlapping framing windows. Causal 1D convolutions extract local features without future lookahead. Rolling buffer retaining K previous step embeddings for temporal context. Cross-attends current frame queries against cached past keys/values to refine features. Reconstructed discrete speech sources separated from background noise.