Attention in Krylov space: Transformer-based extrapolation of Lanczos coefficients
Verdict
该工作将Lanczos系数外推表述为因果序列预测问题,并提出Transformer自回归模型;摘要报告其在混沌、可积及跨系统尺寸任务中优于传统渐近拟合。
论文利用Transformer从短Lanczos系数前缀自回归预测深层Krylov空间中的后续系数,从而以低于直接Lanczos迭代的成本重构算符动力学相关可观测量。
研究问题
在直接Lanczos迭代受到数值不稳定性和内存成本限制时,能否用学习系数历史结构的序列模型,比渐近形式拟合更准确地外推Lanczos系数并重构算符动力学?
方法线索
- 将Lanczos系数视为具有因果结构的时间序列。
- 构建Transformer模型,根据短系数前缀自回归预测未来Lanczos系数。
- 在经典和量子混沌系统中比较Transformer外推与传统渐近拟合。
- 评估外推系数对物理可观测量重构的效果。
- 测试模型在可积区间中的系数外推能力。
- 开展跨系统尺寸迁移:在较小系统训练后,用于较大系统而不再训练。
- 通过注意力模式分析与目标注意力消融,识别对预测最重要的系数历史区段。
对你的用途
- Krylov复杂度、算符增长和Lanczos系数的长程数值外推。
- 在内存或数值稳定性受限时重构深层Krylov空间动力学。
- 比较混沌与可积系统中的Lanczos系数结构。
- 将序列建模方法用于量子多体动力学计算的代理模型设计。
- 分析哪些早期Lanczos系数历史对远期预测最关键。
可核查证据
Abstract摘要指出,直接计算Lanczos系数会受到数值不稳定性和内存成本限制,而传统早期系数渐近拟合可能遗漏影响可观测量重构的次级历史结构。
Abstract作者提出Transformer自回归模型,以短Lanczos系数前缀预测后续系数。
Abstract摘要报告该模型在经典与量子混沌系统中优于渐近拟合,并将误差降低约一个数量级;模型还可处理无通用渐近拟合形式的可积区间。
Abstract摘要称模型可从小系统训练迁移至大系统外推,且无需重新训练。
量化结果
- 摘要报告:相对于渐近拟合,Transformer在Lanczos系数外推和物理可观测量重构中实现了约一个数量级的误差降低。
仍需核实
- 摘要未说明Lanczos系数对应何种初始算符、内积定义或具体动力学生成元。
- 摘要未给出Transformer的层数、参数量、上下文窗口、优化器或训练数据生成方式。
- “约一个数量级”的误差改进缺少具体指标、数值区间和不同任务间的分项结果。
- 摘要未明确物理可观测量的种类,以及它们由外推系数重构的具体方法。
- 无法仅凭摘要判断模型是否保持Lanczos系数的正性、渐近约束或其他物理一致性条件。
展开原始摘要
arXiv:2601.07937v2 Announce Type: replace-cross Abstract: The Universal Operator Growth Hypothesis formulates time evolution of operators through Lanczos coefficients. In practice, however, numerical instability and memory cost limit the number of coefficients that can be exactly computed. In response to these challenges, the standard approach relies on fitting early coefficients to asymptotic forms, but such procedures can miss subleading, history-dependent structures in the coefficients that subsequently affect reconstructed observables. In this work, we treat the Lanczos coefficients as a causal time sequence and introduce a transformer-based model to autoregressively predict future Lanczos coefficients from short prefixes. For classical and quantum chaotic systems, our model outperforms asymptotic fits in both coefficient extrapolation and physical observable reconstruction, and achieves an order-of-magnitude reduction in error. The model also accurately extrapolates coefficients in integrable regimes, where no universal asymptotic fit exists. Remarkably, our model transfers across system sizes: it can be trained on smaller systems and then be used to extrapolate coefficients on a larger system \emph{without retraining}. By probing the learned attention patterns and performing targeted attention ablations, we identify portions of the coefficient history that are most influential for accurate forecasts. Our results demonstrate that modern sequence models can serve as practical surrogates for probing operator dynamics deep in Krylov space, where brute-force Lanczos iteration can be computationally prohibitive.