Skip to content

Transformer 面试必考:自注意力机制详细拆解 + 为什么它比 RNN 更适合长序列?

2260 字约 8 分钟

transformerattentioninterview

2026-06-16