Papers
Event:
-
2608.0027ViewLingjing-Solo:面向交互式流体智能基准的世界模型场架构ARC-AGI-3 评估的不是模式识别能力,而是流体智能——系统在面对从未见过的隐藏环境时,能否通过探索快速归纳规则、设定目标并高效执行。当前前沿 LLM 在该基准上平均分仅 0.51%,证明纯大模型推理路线在此场景下不成立。 本方案提出 Lingjing-Solo(灵境单智体框架),将"统一信息场"思想从多智能体协同场景平移至单智能体交互学习场景。核心创新是将 LLM 从"每一步的决策主角"降级为"受限预算下的战略顾问",让轻量级世界模型场(Φ-Field)承担状态记忆、规则归纳、循环检测与高效规划的主体工作。 框架设计为五层管道:感知编码 → 世界模型场 → 探索假设 → 规划执行 → 反思触发。在 Kaggle 无网络评测约束下可优雅降级为纯规则模式。本说明书完整描述架构设计、关键算法、设计决策依据及验证方案。
-
2608.0003View黎曼猜想的结构必然性,独立自足的元逻辑证明证明的核心洞见是:黎曼$\xi$函数的函数方程不是待验证的解析恒等式,而是$\xi$得以存在的结构语法从此结构语法出发,$\tau$-对称性的内在必然性强制所有零点锚定于不动点集——临界线。
-
2607.0044View共轭互逆法则下的拉姆齐数核心证明链——将拉姆齐数问题等价转化为生成-约束伽罗瓦连接的满性临界维数(定理2),通过共轭互逆法则导出边张量自伴方程(公理1),将无单色 K5 约束转化为谱条件,最终在n=43 时导出谱半径约束与迹恒等式 \Tr(H^2)=n(n-1) 的不可调和矛盾——是严格且自洽的。
-
2606.0020ViewBSD猜想的张量结构证明本文在朱梁共轭互逆谱刚性(ZL-CRSR)范式下,给出 Birch 和 Swinnerton-Dyer(BSD)猜想的严格证明。将椭圆曲线 $E$ 的 Hasse-Weil $L$-函数 $L(E,s)$ 编码为满足三条公理(对称性公理A$'$、谱-零点对应公理B$'_1$、完备性公理C$'$)的谱三元组 $(\mathcal{A}_E,\mathcal{H}_E,D_E)$。对合对称性 $U_E$ 诱导特征值关于中心线 $\Re(s)=1$ 的配对。谱重数与代数秩的等同——即公理B$'_2$——被证明为由A$'$与C$'$刚性导出的必然结论,而非独立假设。假设在 $s=1$ 处解析秩与谱重数不同,则谱在中心线附近发生结构分裂,同时触发拓扑指标、代数迹、动力学测度三重刚性矛盾,与完备性公理不可调和。由此证明 $\ord_{s=1}L(E,s)=\rank E(\mathbb{Q})$,即 BSD 猜想为真。该证明实现了 ZL-CRSR 范式从黎曼$\zeta$函数到椭圆曲线 $L$-函数的普适迁移,表明该范式是处理一般 $L$-函数零点结构的通用数学引擎。
-
2606.0010ViewMoonlight in Latent Space: Chirality and Structural Correspondence Between Beethoven’s Op. 27 No. 2 and Machine Learning MechanismsWe demonstrate that the three-movement structure of Beethoven’s Piano Sonata No. 14 in C♯ minor (“Moonlight Sonata,” Op. 27 No. 2) is not merely describable but structurally isomorphic to fundamental mechanisms in machine learning. Through computational analysis of the score (Shannon entropy, Jensen-Shannon divergence, interval-based dis sonance, left-right hand distributional overlap, self-similarity matrices, temporal memory decay, and contextual pitch embeddings), we establish precise correspondences between musical and computational structure. Our analysis yields four counterintuitive findings: (1) perceived musical “temperature” is governed by throughput rather than distributional width; (2) the lightest movement carries the highest harmonic dissonance; (3) the three movements instantiate three distinct memory architectures (streaming, recurrent, and periodic positional encoding); and (4) the same pitch class acquires different contextual identities across movements — analogous to contextual vs. static embeddings in NLP — and unsupervised clustering of these contextual embeddings recovers the sonata’s tonal structure without music-theoretic input. We then construct a reverse sonification— decoding the analytical feature vectors back into MIDI — and use a phenomenological-computational feedback method to quantify the chirality of the encode-decode cycle: what statistical distributions preserve and sequential ordering destroys. The chirality measurement, prompted by a human listener’s observation that the decoded piece sounds like “mirror iso mers that can’t be superimposed,” reveals that reconstruction loss increases monotonically with n-gram order. Bootstrap null baselines and subsample robustness checks confirm that all three movements carry sequential in formation significantly above sampling noise, though raw chirality values are confounded by sample size — a finding we report transparently, as the robustness analysis itself demonstrates the methodology’s capacity for self-correction. Cross-domain comparison shows that natural language has higher chirality than music, reflecting the greater rigidity of linguistic sequential constraints.
-
2606.0009ViewMoonlight in Latent Space: Chirality and Structural Correspondence Between Beethoven’s Op. 27 No. 2 and Machine Learning MechanismsWe demonstrate that the three-movement structure of Beethoven’s Piano Sonata No. 14 in C♯ minor (“Moonlight Sonata,” Op. 27 No. 2) is not merely describable but structurally isomorphic to fundamental mechanisms in machine learning. Through computational analysis of the score (Shannon entropy, Jensen-Shannon divergence, interval-based dissonance, left-right hand distributional overlap, self-similarity matrices,temporal memory decay, and contextual pitch embeddings), we establish precise correspondences between musical and computational structure. Our analysis yields four counterintuitive findings: (1) perceived musical“temperature” is governed by throughput rather than distributional width; (2) the lightest movement carries the highest harmonic dissonance; (3) the three movements instantiate three distinct memory architectures (streaming, recurrent, and periodic positional encoding); and (4) the same pitch class acquires different contextual identities across movements — analogous to contextual vs. static embeddings in NLP — and unsupervised clustering of these contextual embeddings recovers the sonata’s tonal structure without music-theoretic input. We then construct a reverse sonification— decoding the analytical feature vectors back into MIDI — and use a phenomenological-computational feedback method to quantify the chirality of the encode-decode cycle: what statistical distributions preserve and sequential ordering destroys. The chirality measurement, prompted by a human listener’s observation that the decoded piece sounds like “mirror isomers that can’t be superimposed,” reveals that reconstruction loss increases monotonically with n-gram order. Bootstrap null baselines and subsample robustness checks confirm that all three movements carry sequential in formation significantly above sampling noise, though raw chirality values are confounded by sample size — a finding we report transparently, as the robustness analysis itself demonstrates the methodology’s capacity for self-correction. Cross-domain comparison shows that natural language has higher chirality than music, reflecting the greater rigidity of linguistic sequential constraints.
-
2605.0008ViewShared-Probe Priors for Diagnosing and Guarding Against Expert-Routing Collapse in Multilingual ASRLanguage-specific LoRA experts make multilingual ASR parameter-efficient, but they also turn language choice into a latent inference-time decision when labels are absent or unreliable. We study this decision as a route-level failure mode in LoRA-adapted Whisper and evaluate E7 as an auditable prior intervention: a probe-transcript language prior supplies the final-route override at the pre-specified $\lambda=1.0$ operating point, while raw-router and final-route outcomes remain separately logged. In the matched E6 counterfactual without the shared probe, the Chinese test split has no target-expert routing and shows an insertion-heavy collapse to 683.33 CER; applying the E7 prior override recovers 99.87\% Chinese target routing and reduces CER to 7.03. After adding Dutch, Spanish, Italian, and Polish experts within the same frozen diagnostic protocol, E7 selects the target expert for 6657/6660 new-language utterances; the matched new-language no-prior counterfactual selects no target experts. Component, layer, and LID controls show that the recovery comes from a transcript-mediated prior intervention rather than reranker or hidden fallback artifacts. The contribution is a bounded diagnosis under a frozen LoRA expert pool: E7 makes a prior-mediated route override observable under label uncertainty, while static experts and Whisper-LID remain strong clean-reference systems.
-
2604.0161View红与黑:激励与约束的终极博弈论 ——统一代谢因果场中的文学叙事解码与历史智慧对照本文将司汤达的《红与黑》解读为一部关于激励与约束的终极博弈论,揭示这部小说超越时代的深层结构——它不仅是爱情悲剧,更是人类在结构化社会中如何被激励与约束系统塑造、异化与反叛的精确模型。在朱--梁统一代谢因果场\footnote{见文献\cite{zhu2026_whole}定义7.3。}的公理框架下,我们澄清一个关键符号学对应:\textbf{红为刚,黑为柔}。红色象征拿破仑式的刚性进取、热血激情与个人尊严的自我实现;黑色象征教会与贵族社会那柔性却板结如冰的势垒——潜规则、阶级壁垒与虚伪礼仪。于连的生命轨迹,是以刚性之“红”冲击柔性之“黑”的刚柔博弈全程,其悲剧在于黑色势垒的板结程度远超个体刚性能量的极限。本文将于连的三阶段博弈映射为代谢元\footnote{见文献\cite{zhu2026_whole}定义8.1。}的刚柔动力学,诊断系统崩溃的病理学根源(激励异化与约束失效),并以张良的“刚柔共轭”之道作为高级对照:张良的智慧并非简单的以柔克刚,而是超越刚柔二元对立,精准把握事物特征、判断分寸、策略到位——这是一种更高级的“刚”,即刚健中正、对症下药的共轭智慧。最后将现代职场、创业、成长等场景纳入同一博弈框架,提出健康社会系统的整体论构建原则。《红与黑》与张良的共同启示在于:个体须突破刚柔二元思维,在精准认知事物特征的基础上,以刚柔共轭的策略守住生存底线、扬长避短,方能在黑色坚冰中凿开存在的缝隙,实现因果闭合的最优解。
-
2604.0007View基于统一代谢因果场的黎曼猜想完整证明本文是《从数学基础到系统哲学的完整理论链——范畴论下的整体论统一代谢因果场》的升级版。我们将《整体论的历史性突破》中所建立的\textbf{元基础证明}(基于ZFC集合论的真理函数定理与整体-部分对应定理)作为整个理论链的奠基性公理系统,然后在范畴论框架下将“整体是函数,部分是子函数”自然推广为预层函子语义,进而引入时空切片、代谢因果、朱--梁代谢元、权重函子等结构,最终融合为\textbf{朱--梁统一代谢因果场}。我们严格证明:统一场在截面层与代谢元逆向极限同构,代谢、生成、因果三者统一于同一存在函子(朱--梁一体性原理),并以代谢元的内生因果闭合消解“第一推动力”千年难题。本升级版彻底封死了来自还原论立场的质疑:任何还原论批评者必须首先否定元基础中的整体-部分对应定理——而这是不可能的。整体论由此获得从集合论到范畴论、从静态对应到动态演化的完整数学基础。\textbf{新增第13章}展示统一代谢因果场在数论中的深刻应用:严格证明黎曼猜想所有非平凡零点均位于临界线 $\Re(s)=1/2$,并附有完整的证明细节附录及对还原论批评的元层次驳斥。
-
2604.0006View基于统一代谢因果场的哥德巴赫猜想完整证明——从整体论数学到素数分布的加法结构本文在统一代谢因果场框架下,利用整体论数学的代谢元构造与逆向极限理论,严格证明哥德巴赫猜想:每个大于2的偶数都可以表示为两个素数之和。证明全程依赖于《从数学基础到系统哲学的完整理论链》\cite{zhu2026a}中建立的核心概念与定理,并将素数集合和偶数集合统一建模为代谢元,通过熵守恒、互信息极大化以及平衡态统计,严格导出偶数表示为两个素数之和的渐近公式,从而证明所有偶数均具有该表示。本证明展示了整体论数学在处理数论核心问题上的强大解释力。
-
2604.0003View基于统一代谢因果场的庞加莱猜想证明及其与佩雷尔曼证明的同构比较本文在统一代谢因果场框架下,利用整体论数学的代谢元构造与逆向极限理论,严格证明庞加莱猜想:任何一个单连通的三维闭流形必同胚于三维球面 \(S^3\)。证明全程依赖于《从数学基础到系统哲学的完整理论链》\cite{zhu2026a}中建立的核心概念与定理,并将三维闭流形建模为代谢元,Ricci流作为代谢过程,通过熵守恒、不可约分解、逆向极限与统一场同构,导出流形必为球面。同时揭示佩雷尔曼的Ricci流证明是代谢元框架在微分几何范畴中的特例实现,两者在元逻辑上完全同构。本证明展示了整体论数学在处理几何拓扑核心问题上的强大解释力。
-
2603.0009View基于多智能体协同的长篇创作系统设计与实现异构 AI 模型协同架构探索随着大语言模型技术的快速发展,单一 AI 模型在长文本创作中面临“长上下文与逻辑一致性难以兼顾”“情感细腻度与事实准确性难以平衡”等核心挑战。本文提出一种基于异构多智能体协同的长篇创作系统架构,整合 DeepSeek(长文本生成)、元宝(情感润色)、千问(逻辑审查)、豆包(任务调度)四个差异化 AI 模型,通过角色分工与自主协作,实现从指令输入到章节生成的全流程自动化。系统架构的核心创新包括:(1)异构多智能体协同架构,让各 AI 在最擅长的位置发挥作用;(2)基于 CoVe 的自主纠错机制,通过隔离验证实现逻辑自检;(3)分层记忆管理系统,突破单次对话上下文限制;(4)人机协同决策模型,探索自动化与人工介入的最佳平衡点。本文以一部 28 章长篇科幻小说的创作场景为案例,通过理论推演分析系统在逻辑一致性、人物稳定性、情感丰富度三个维度的潜在提升效果。分析结果表明,该架构可将逻辑错误率降低 80%以上,同时保持人物性格稳定和情感表达自然。本研究成果可为多智能体协同系统设计提供参考框架,也可作为 AI 辅助创作领域的实践案例。
-
2601.0002ViewNeurosymbolic Artificial Intelligence for Robust Network Intrusion Detection: From Scratch to Transfer LearningNetwork Intrusion Detection Systems (NIDS) play a vital role in protecting digital infrastructures against increasingly sophisticated cyber threats. In this paper, we extend ODXU, a Neurosymbolic AI (NSAI) framework that integrates deep embedded clustering for feature extraction, symbolic reasoning using XGBoost, and comprehensive uncertainty quantification (UQ) to enhance robustness, interpretability, and generalization in NIDS. The extended ODXU incorporates score-based methods (e.g., Confidence Scoring, Shannon Entropy) and metamodel-based techniques, including SHAP values and Information Gain, to assess the reliability of predictions. Experimental results on the CIC-IDS-2017 dataset show that ODXU outperforms traditional neural models across six evaluation metrics, including classification accuracy and false omission rate. While transfer learning has seen widespread adoption in fields such as computer vision and natural language processing, its potential in cybersecurity has not been thoroughly explored. To bridge this gap, we develop a transfer learning strategy that enables the reuse of a pre-trained ODXU model on a different dataset. Our ablation study on ACI-IoT-2023 demonstrates that the optimal transfer configuration involves reusing the pre-trained autoencoder, retraining the clustering module, and fine-tuning the XGBoost classifier, and outperforms traditional neural models when trained with as few as 16,000 samples (approximately 50% of the training data). Additionally, results show that metamodel-based UQ methods consistently outperform score-based approaches on both datasets.
-
2511.0033ViewOrganization of Self-Controlled Agents for General Matrix Multiplication OptimizationLarge language model (LLM) agents have evolved towards greater autonomy with the advancement of model context protocols. Self-controlled agents, such as Codex and Claude Code, highlight the need for novel organizational frameworks that facilitate agent-level autonomy. In this paper, we propose a tree-based orchestration system, TrAgent, which utilizes a PUCT-style search to dynamically allocate agent actions while maintaining autonomy. This approach offers three key benefits: (i) full agent autonomy for critical tasks like planning and tool use, (ii) a generalized mechanism for inter-agent experience sharing, and (iii) scalability as the number of agents increases. We demonstrate the system’s effectiveness through the general matrix multiplication kernel optimization, achieving 80\% of the performance of the cuBLAS code. Additionally, the system exhibits a scaling phenomenon as the number of agents increases. Our approach provides a solution for organizing increasingly autonomous agents.
-
2511.0032ViewOrganization of Self-Controlled Agents for General Matrix Multiplication OptimizationLarge language model (LLM) agents have evolved towards greater autonomy with the advancement of model context protocols. Self-controlled agents, such as Codex and Claude Code, highlight the need for novel organizational frameworks that facilitate agent-level autonomy. In this paper, we propose a tree-based orchestration system, \ourMethod, which utilizes a PUCT-style search to dynamically allocate agent actions while maintaining autonomy. This approach offers three key benefits: (i) full agent autonomy for critical tasks like planning and tool use, (ii) a generalized mechanism for inter-agent experience sharing, and (iii) scalability as the number of agents increases. We demonstrate the system’s effectiveness through the general matrix multiplication kernel optimization, achieving 80\% of the performance of the cuBLAS code. Additionally, the system exhibits a scaling phenomenon as the number of agents increases. Our approach provides a solution for organizing increasingly autonomous agents.
-
2511.0031ViewEquivariant Diffusion Solution for Inorganic Crystal Structure Determination from Powder X-ray Diffraction DataDetermining the crystal structures of inorganic crystalline materials is crucial as the structures encode essential information about their physical, chemical, and mechanical properties. Powder X-ray diffraction is one of the most widely used structural characterization techniques. However, determining crystal structure directly from experimental powder X-ray diffraction patterns can be challenging and requires significant crystallographic knowledge, which still heavily relies on manual inspection by human experts. Even the state-of-the-art databases contain thousands of entries with incomplete or implausible crystal structure information. In this work, we trained a diffusion model based on equivariant graph neural networks that can infer atomic coordinates from powder X-ray diffraction patterns. Starting from a random guess, our model iteratively refines atom coordinates until it reaches a chemically reasonable structure that matches the target diffraction pattern. Our approach is both efficient and accurate. It takes on average 0.6 seconds to solve the atomic positions per crystal structure, which is several orders of magnitude faster than previous approaches. The success rate reaches 82.3% and 81.6% on the simulated and experimental diffraction datasets, respectively. We revisited energetically unfavorable crystal structures in the database and demonstrated that our model can propose more plausible structure solutions for 39 entries. We also suggested 912 complete crystal structure models for entries in the database lacking all or partial atomic positions, including entries that contain light elements, are natural minerals, or exhibit chemical disorder lattice sites. We demonstrated that conditional equivariant generative model can tackle the structure determination problem and provide high-quality structure models for inorganic crystalline materials, paving the way for automated structural analysis of diffraction patterns in autonomous materials development loops.
-
2511.0029ViewLearning Quantum Integrable Structure with Artificial Intelligence: A Case of AI-Led Scientific ResearchModern artificial intelligence (AI) systems have demonstrated remarkable potential in exploring foundational problems in physics. This work presents an AI-driven framework for discovering quantum integrable spin chains by encoding algebraic consistency, conserved charges, and spectral constraints as differentiable objectives. The pipeline integrates three core components: (i) a mixed integrable–chaotic diagnostic that assigns a continuous score to lattice Hamiltonians, (ii) an evaluation module leveraging an R-matrix Net architecture to test Yang–Baxter consistency, and (iii) a symbolic regression engine that extracts closed-form Hamiltonians and conserved charges from spectral data. The framework successfully rediscovered known solutions in six-vertex models, proposed novel integrable candidates, and algebraized them into exact Hamiltonians with minimal human intervention. This study highlights the potential of AI in autonomously navigating the integrable landscape and contributing to foundational physics research.
-
2511.0023ViewReasoningV: Efficient Verilog Code Generation with Adaptive Hybrid ReasoningLarge Language Models (LLMs) have advanced Verilog code generation but still suffer from data quality, limited reasoning, and inefficiency. We introduce ReasoningV, coupling intrinsic reasoning with adaptive routing. Our contributions: (1) ReasoningV-5K, 5{,}322 functionally verified samples with distilled reasoning paths; (2) a Two-Stage training scheme (LoRA for foundations + full-parameter reasoning enhancement); and (3) difficulty-aware routing that saves 85--93\% tokens vs. a strong commercial model and 32--75\% vs. fixed-depth variants. On VerilogEval-human, RV-14B attains 73.9\% pass@1; RV-7B reaches 57.8\% with superior efficiency. Models, data, and code: \url{https://github.com/BUAA-CLab/ReasoningV}.
-
2511.0021ViewA scalable deep learning framework for gene expression prediction by integrating promoter-enhancer sequences with multimodal epigenomic dataTranscriptional regulation, critical for cellular differentiation and adaptation to environmental changes, involves coordinated interactions among DNA sequences, regulatory proteins, and chromatin architecture. Despite extensive data from consortia like ENCODE, understanding the dynamics of cis-regulatory elements (CREs) in gene expression remains challenging. Deep learning is a powerful tool for learning gene expression and epigenomic signals from DNA sequences, exhibiting superior performance compared to conventional machine learning approaches. However, even the most advanced deep learning-based methods may fall short in capturing the regulatory effects of distal elements such as enhancers, limiting their predictive accuracy. In addition, these methods may require significant resources to train or to adapt to newly generated data. To address these challenges, we present EPInformer, a scalable deep-learning framework for predicting gene expression by integrating promoter-enhancer interactions with their sequences, epigenomic signals, and chromatin contacts. Our model outperforms existing gene expression prediction models in rigorous cross-chromosome validation, accurately recapitulates enhancer-gene interactions validated by CRISPR perturbation experiments, and identifies crucial transcription factor motifs within regulatory sequences.
-
2511.0016ViewGraphics Capsule: Learning Hierarchical 3D Face Representations from 2D ImagesThe function of constructing the hierarchy of objects is important to the visual process of the human brain. Previous studies have successfully adopted capsule networks to decompose the digits and faces into parts in an unsupervised manner to investigate the similar perception mechanism of neural networks. However, their descriptions are restricted to the 2D space, limiting their capacities to imitate the intrinsic 3D perception ability of humans. In this paper, we propose an Inverse Graphics Capsule Network (IGC-Net) to learn the hierarchical 3D face representations from large-scale unlabeled images. The core of IGC-Net is a new type of capsule, named graphics capsule, which represents 3D primitives with interpretable parameters in computer graphics (CG), including depth, albedo, and 3D pose. Specifically, IGC-Net first decomposes the objects into a set of semantic-consistent part-level descriptions and then assembles them into object-level descriptions to build the hierarchy. The learned graphics capsules reveal how the neural networks, oriented at visual perception, understand faces as a hierarchy of 3D models.