Papers
Event:
-
2510.0052ViewConformal Prediction as Bayesian Quadrature for Risk ControlIn this paper, we present a novel framework that leverages Bayesian quadrature for conformal prediction to achieve rigorous, data-conditional, and distribution-free risk guarantees, addressing the challenge of controlling predictive risk in high-stakes, black-box settings. Our approach constructs an upper bound on the expected loss by integrating over the quantile function of the loss distribu- tion, where, given calibration losses ℓ1 , . . . , ℓn , we define the aggregated loss Pn+1 as L+ = i=1 Ui ℓ(i) with Dirichlet random variables Ui ∼ Dir(1, . . . , 1) and ℓ(n+1) = B, thereby ensuring that the condition Pr(L+ ≤ α) ≥ β is met. Our contributions include a principled derivation that recovers well-known conformal methods such as Split Conformal Prediction (SCP) and Conformal Risk Control (CRC) as special cases, while introducing a novel high posterior density (HPD) rule that exploits the full posterior of L+ . We rigorously validate our method on synthetic binomial loss and heteroskedastic regression tasks, where experimental results indicate that methods based solely on the posterior mean (CRC) or uniform concentration bounds (RCPS) often yield either overly optimistic or conservative decisions, whereas our HPD rule achieves risk control with zero empirical failure rate and improved utility. For example, in the binomial experiment, while SCP selects an average λ of 0.596 with a 61.6% failure rate, HPD selects λ ≈ 0.970 with a 0% failure rate, and a similar trend is observed in regression tasks with test risks decreasing from 0.512 for SCP to 0.067 for HPD. These findings, summarized in Table 1, confirm that our Bayesian quadrature reformulation not only provides a more interpretable statistical characterization of conformal risk but also adapts effectively to calibration sample size and confidence level tuning, thus offering a robust solution for high-stakes decision-making.
-
2510.0051ViewCOMD: Coherent Masked DiffusionMasked language models (MLMs) have shown promise in natural language processing, but struggle with generating coherent and coherent-sounding text. In this work, we present Coherent Masked Diffusion (CoMD), a novel framework that extends Masked Language Diffusion to more efficiently and more effectively learn coherent and incoherent language. CoMD is built on Masked Language Diffusion (MLD), a recently proposed framework that models text generation as an inverse denoising diffusion process. Unlike MLD, CoMD uses a fixed mask matrix that is independent of the masked-out token and optimizes the probability of coherent generations with a novel coherent loss term without requiring additional samples per training step. Additionally, CoMD uses a variable time parameter to guide the coherent probability towards the ground truth coherent probability. Both inference and training computation are constant with respect to the length of the text. Empirically, CoMD outperforms previous methods on multiple coherent benchmarks. Furthermore, CoMD achieves an inference speedup of 7.3x and 10.5x over MLD and MDLM, respectively, and is significantly more compute and parameter efficient than autoregressive models.
-
2510.0050ViewU-CAN: User-Guided Clarification for Asking Clarification in Asking Across Needs FrameworkIt is still unclear if and how methods developed specifically on asking clarification for retrieval or problem-solving in the academic community can effectively address user needs during human-computer interactions (HCI). In this work, we first propose an Asking Across Needs (AAN) framework to explore the complexities of HCI, including user needs, interaction styles, and interaction types, by building an interaction graph (Pearl, 2009) containing user and LLM actions. Then, we create a new benchmark, UsClarification for Asking Needs (U-CAN), containing task-oriented asking clarification and retrieval-related asking clarification which align with real-world HCI scenarios. Specifically, we design new interaction graph designs and user-guided prompting techniques based on our AAN framework to address multiple user needs not met in existing HCI studies. We find that task-oriented needs are often left unmet, and existing methods show performance gaps between simulated and real-world (enrolled students) settings. We also demonstrate that HCI can be facilitated by interaction graphs on retrieval-related asking clarification using our proposed interactive graph model.
-
2510.0049ViewLearning Unnormalized Models with Missing Data via Adversarial Score MatchingLearning unnormalized model parameters is a challenging task that frequently arises in various scientific fields. Score matching is a promising method to learn unnormalized models by estimating the score function. However, score matching has several practical challenges in real-world applications, including the need for an auxiliary network to estimate the score function, the requirement for the model to support sampling, and the difficulty of estimating the score function for high-dimensional data. To address these challenges, we propose adversarial score matching (ASM), an adversarial learning algorithm for learning unnormalized models, which does not require an auxiliary network and can be applied to high-dimensional data. We also propose a multilevel Monte Carlo estimator for the score discrepancy, which is computationally more efficient than the traditional importance sampling estimator. In addition, we demonstrate that ASM is a mode-seeking algorithm, which has been observed empirically in a variety of adversarial learning methods. We evaluate the performance of ASM on various unnormalized models and missing data mechanisms, and demonstrate that ASM outperforms existing score matching methods.
-
2510.0048ViewRisk Control With Width-Sketching AlgorithmsWe introduce the notion of width-sketching algorithms, defined as algorithms with provably bounded width (that is, probability of containing the randomness) for the induced coverage set. For algorithms that sketch the width, we prove a novel uniform upper bound and provide an instance where the width in expectation is twice as large as the optimal width. We then introduce the width-optimality notion and an approximate version termed mean-width optimality, which allows us to derive algorithms with the desired coverage while minimizing the mean width. We provide a high-level perspective on the relationship with depth-sketching algorithms, i.e., algorithms that sketch the depth of the induced sets with probability 1 − α, and show that they provide complementary forms of coverage. Finally, we demonstrate the application of the framework to conformal prediction with Bayesian quadrature.
-
2510.0047ViewPredictive Need Assessment, Public Service Providers and Inequalities of Labor Market OutcomesAid and assistance are key to reducing social inequalities. In the public sector, aid providers face the challenge to distribute resources according to needs. In recent years, algorithms for need assessment have become an integral part of public institutions. While need assessment is crucial, the implications for the distribution of aid and assistance in the public sector are still poorly understood. In this work, we investigate how the use of predictive models for need assessment impacts the distribution of aid and assistance in the German public employment service. To this end, we develop a synthetic dataset for the “first round” assignment in the German public employment service based on regional data from the State of Bavaria in 2019. Our dataset comprises labels for 85,299 out of 275,889 employed and unemployed, treated and untargeted individuals in 2019. The label indicates whether a person received a prioritized status and thus privileged access to government services. We find that the use of predictive models leads to significant resource imbalances and deepens the divides along the lines of migration background, education, and gender. These findings highlight important ethical implications for public service providers in need of need assessment for aid and assistance.
-
2510.0046ViewEconomic Implications of Language Models and Copyright LawHow will language models (LMs) affect future economic progress? Inspired by the Lever of Riches by Mokyr (1992), we argue that the institutions governing LM content generation and usage patterns are critical to answering this question. We content that, because LM creators have a strong incentive to collect, train, and deploy intellectual property protection, the all-you-can-consume access to knowledge and creativity they enable has led to rapid acceptance and widespread use, which in turn results in smaller, low-skill employment creation but increased output and greater overall welfare. We provide a theoretical and analytical framework explaining this phenomenon and point to its long-term consequences using empirical evidence.
-
2510.0045ViewPST-AUTO-AGENT: A Multi-Agent Ensemble Framework for Paper Source TracingThe escalating volume of scientific literature necessitates efficient methods for identifying foundational works that significantly inform new research. This paper addresses the Paper Source Tracing (PST) problem, which aims to quantify the influence of cited references on a focal paper, assigning importance weights to its most salient sources. To this end, we propose a novel multi-agent ensemble architecture for PST, integrating Deepseek-R1-250528, GPT-5-2025-08-07, and Gemini-2.5-pro. Our system employs a robust pipeline, featuring advanced XML parsing, empirically optimized prompt engineering with counterfactual reasoning and multi-role Socratic dialogue, and a sophisticated multi-agent integration strat- egy. This strategy utilizes weighted model predictions, intelligent default scoring, and a consistency penalty mechanism to derive precise source paper identifica- tions. Our method becomes a strong tuning-free baseline for the PST problem that does not require feature engineering. Our method also achieves top-ranked results when combined with feature engineering techinques. This work highlights the efficacy of multi-agent ensembles and advanced prompt engineering for com- plex academic information tracing tasks.
-
2510.0044ViewA Comprehensive Survey on Deep LearingMachine learning and deep learning methodologies have revolutionized computational approaches to complex problem-solving across numerous domains, emerging as transformative technologies in artificial intelligence research [1,7]. This comprehensive review synthesizes current literature to examine the theoretical foundations, methodological advancements, and practical implementations of these techniques, highlighting their evolution from basic machine learning concepts to sophisticated deep neural architectures [2,9]. The analysis demonstrates remarkable success in applications ranging from computer vision and natural language processing to healthcare diagnostics and autonomous systems, with deep learning models achieving unprecedented performance in pattern recognition tasks [3,4,8]. However, significant challenges persist, including the need for massive labeled datasets, computational resource requirements, model interpretability issues, and inherent parameter redundancy in deep architectures [5,6]. The review identifies emerging opportunities in transfer learning, few-shot learning, and explainable AI as promising research directions [10]. By critically evaluating both current limitations and future potential, this analysis provides a structured framework for researchers to advance the field while addressing practical implementation barriers across diverse application domains.
-
2510.0043ViewDecoupling Openness and Connectivity: Non-Monotonic Effects in LLM-Based Cultural DynamicsCultural dynamics in multi-agent systems exhibit a counterintuitive phenomenon: local similarity-based interactions can lead to global fragmentation rather than convergence. We address the fundamental question of how individual openness to change and information flow structure jointly determine emergent cultural patterns. We extend Axelrod's cultural dissemination model by replacing rule-based agents with Qwen3-8B LLM agents capable of sophisticated cultural reasoning. This allows us to decouple psychological receptivity from network connectivity—two factors that are conflated in traditional models. Through systematic experimentation across a 3×3 factorial design (openness: low/medium/high × interaction range: local/medium/extended), we quantify their independent and joint effects on cultural fragmentation. Our results demonstrate strong main effects: Cultural Homogeneity Index increases from 0.279 to 0.437 with higher openness (1st order interactions, +57\%), while optimal information flow (3rd order) achieves the highest convergence at 0.489 for high openness agents—representing 75\% improvement over low openness baseline (0.279). Critically, we uncover a non-monotonic relationship where 3rd-order interactions consistently outperform both 1st and 5th-order across all openness levels, revealing an optimal balance between exploration and exploitation. Code can be found at https://anonymous.4open.science/r/YuLan-OneSim/.
-
2510.0037ViewState-Dependent Dynamics Among Apple Stock, Bitcoin, and Gold: Evidence from Rolling Correlations, Connectedness, and Tail CopulasMega-cap technology equities, cryptocurrencies, and gold increasingly co-determine modern portfolio outcomes, yet their joint dynamics are regime-dependent and incompletely understood. This paper studies the state-dependent comovement among Apple Inc. (AAPL), Bitcoin (BTC), and gold (XAU) using daily data from 2015 to 2025. We triangulate three complementary lenses: (i) rolling Pearson correlations to trace smooth co-movement and structural shifts; (ii) a Diebold–Yilmaz connectedness framework based on rolling 252-day VARs and generalized forecast-error variance decompositions to identify directional risk transmitters and receivers; and (iii) empirical copulas to estimate lower- and upper-tail dependence at the 5% threshold, contrasting crash versus rally dynamics. We document three core results. First, the AAPL–BTC correlation rose from near zero before 2020 to roughly 0.30 during the COVID-19 period and has remained persistently elevated through 2025, indicating a lasting post-pandemic regime. Second, gold emerges as a net transmitter of shocks while BTC is a net receiver, revising the canonical view of gold as a purely passive safe haven. Third, AAPL–BTC exhibits pronounced asymmetry in the tails, with downside dependence considerably stronger than upside (“crash correlation”). Robustness checks spanning window sizes, VAR lags, alternative dependence metrics, and tail thresholds corroborate these findings. The portfolio implication is that static diversification across tech, crypto, and gold underperforms precisely when insurance is most needed. We advocate regime-aware allocation that monitors connectedness and transmitter identity, budgets explicitly for joint-tail risk, and uses dynamic overlays. The results also inform macro-prudential monitoring: supervisors should track transmitter rotations and connectedness spikes that presage cross-market stress transmission.
-
2510.0033ViewAI Transformation in Biomedical Research: From Data-Driven to Insight-Driven ApproachesThis review examines the ongoing transformation of artificial intelligence applications in biomedical research, tracing the evolution from data-driven to insightdriven approaches. It synthesizes advances in AI-powered multimodal data integration techniques, including early, intermediate, late, and hybrid fusion strategies that effectively combine heterogeneous biomedical data sources. The review explores how network-based computational frameworks and single-cell technologies are revolutionizing disease mechanism analysis through multi-omics integration, enabling the identification of dysregulated pathways and potential therapeutic targets. It further evaluates AI’s role in enabling precision medicine through personalized diagnostics, treatment selection, and radiomics-based healthcare. The integration of AI with various omics disciplines has enhanced understanding of disease mechanisms at molecular, cellular, and tissue levels, creating unprecedented opportunities for early diagnosis and targeted therapeutics. The review concludes by addressing critical challenges including model explainability and data privacy considerations, while highlighting the emergence of closed-loop AI systems that actively participate in scientific discovery through continuous learning and adaptation. These developments collectively signal a paradigm shift toward AI systems that not only analyze biomedical data but generate actionable insights that advance clinical practice and scientific understanding
-
2510.0032ViewArtificial Intelligence in Biomedical Research: From Data Integration to Precision MedicineThis comprehensive review examines the transformative role of artificial intelli- gence in biomedical research, from foundational data integration to clinical ap- plications. The paper explores how AI techniques facilitate multimodal data fu- sion across diverse biological data types, employing both traditional statistical methods and advanced deep learning architectures including variational autoen- coders, graph neural networks, and transformer models. It evaluates AI appli- cations in medical imaging, where convolutional neural networks have achieved remarkable diagnostic accuracy (up to 94% in COVID-19 detection) while en- hancing segmentation and classification tasks across multiple imaging modalities. The review further investigates generative AI’s impact on molecular design and drug discovery, highlighting transformer-based architectures like TransAntivirus that navigate vast chemical spaces to optimize therapeutic candidates. Finally, it examines AI-enabled precision medicine applications, including Clinical Deci- sion Support Systems and federated learning approaches that balance analytical power with privacy preservation. Despite significant progress, implementation challenges persist, including data heterogeneity, model explainability, and ethical concerns regarding bias and privacy. The paper underscores the importance of developing interpretable AI systems that integrate seamlessly into clinical workflows while addressing regulatory, ethical, and economic considerations to realize the full potential of AI in advancing biomedical research and healthcare delivery.
-
2510.0031View模拟、影响与驯化:受众智能体在新闻传播中的伦理风险与规制路径研究随着生成式人工智能与智能体(Agent)技术的迅猛发展,新闻传播领域正经历从"内容数字化" 向"认知智能化"的范式转型。受众智能体作为能够模拟、预测甚至替代部分人类受众认知与行为的新型数 字实体,其在新闻生产、分发与反馈各环节的深度嵌入,在提升传播效率的同时也引发了复杂的伦理挑战。 本文结合2025年斯坦福大学AI行为研究、中国AI大模型测评报告等最新实证数据,系统审视受众智能体 在新闻传播中的应用所衍生的伦理风险,并构建相应的规制路径。研究发现,受众智能体的伦理风险主要 集中在三个层面:在模拟层面,存在"数字孪生"失真、归因悖论与信任赤字的风险;在影响层面,面临商 业价值侵蚀公共属性、人机协同失当导致价值偏移的困境;在驯化层面,则遭遇技术依赖导致的主体性消 解与规则滞后带来的治理真空。针对上述风险,本文借鉴动态能力理论,提出一个以"感知-捕捉-重构"为 核心的多维治理框架,为新型主流媒体在智能时代的稳健变革提供兼具学理与实践价值的方案。
-
2510.0030ViewLatent-Diffusion Guided Cross-View Alignment for Heterogeneous Graph RecommendationRecommender systems operating on heterogeneous, multi-relational graphs contend with noise and incompleteness in auxiliary signals, which can destabilize learning and degrade ranking performance when targeting robust representations. Naive cross-view training risks propagating noise across views, and existing contrastive or augmentation-based schemes often hinge on design choices and can struggle to scale to large, complex graphs. We propose a latent-diffusion guided cross-view alignment framework for heterogeneous graph recommendation that jointly learns a relation-aware heterogeneous GNN encoder, producing paired target and auxiliary embeddings, and a compact, time-conditioned latent-space denoiser that maps noisy auxiliary latents toward target-view semantics. The denoiser provides principled supervision to disentangle structured noise, with its residual outputs fused into target embeddings to refine ranking-relevant representations. Training optimizes a joint denoising objective and a ranking objective, enabling scalable, robust cross-view alignment without ad-hoc augmentations. Empirical results on implicit-feedback data demonstrate improved robustness and ranking accuracy under noisy auxiliary signals, with flexible gradient-flow and fusion strategies supporting stable end-to-end training on large graphs. Ablations highlight the benefits of explicit noise modeling in auxiliary views, diffusion-based supervision for stability, and scalable, relation-aware encoding of practical significance for recommender systems.
-
2510.0029ViewAI有意识吗?——AI意识的多层次评估框架本文探讨AI是否具有意识这一前沿问题。通过建立一套评估体系,收集整理最新研究结果,对AI的意识水平进行打分评估。基于哲学、神经科学和心理学三个维度的综合分析,结果显示当前AI意识的整体支持度约为43.84%。直观的结果图表可访问 acw.gixia.org 查看。
-
2510.0028ViewEstimating Rural Rooftop Solar Potential Using Semantic Segmentation and Multi-Source DataAbstract. Solar energy, as a clean and renewable resource, has gained significant global attention. In contrast to urban areas, where buildings vary in height and are often obstructed, the relatively flat ru-ral buildings in northern China provide optimal conditions for solar panel installation. Consequently, the solar energy potential of northern rural areas has attracted significant attention from researchers. Traditional studies typically rely on solar radiation simulation software and 3D models to estimate solar radiation and the solar energy potential of buildings. However, the lack of comprehensive and accurate 3D building model data for rural areas in China has significantly hindered progress in this field. To address this limitation, this study proposes a novel method for rapidly estimating the solar energy potential of rural buildings by integrating deep learning algorithms with parametric modeling platforms. Using convolution neural networks (CNNs), the proposed method efficiently and accurate-ly extracts building footprints from complex satellite imagery. These footprints are then imported in-to the Grasshopper parametric platform to generate and optimize vector outlines of buildings. By combining these outlines with digital surface model (DSM) data containing building height infor-mation, the study constructs precise 3D building models. Furthermore, GPU-accelerated solar simula-tion software, Vitality 2.0, is used for rapid solar energy potential estimation. The study conducted building roof extraction based on satellite imagery for 31 villages in Tianjin and generated parametric three-dimensional village models. Through simulation, the research found that due to the relatively low height of village buildings and the absence of mutual shading between buildings, the larger the village scale, the greater the roof area, and consequently, the higher the photovoltaic power genera-tion capacity of the village. The study also revealed that metal roofs, which have better heat dissipa-tion, result in higher photovoltaic panel conversion efficiency. Therefore, compared to villages with roofs primarily made of concrete and ceramic tiles, villages dominated by metal roofs can recoup all the costs of photovoltaic panels in a shorter period.
-
2510.0026ViewGeometry-Aware Optimal Flow Matching via Convex PotentialsGenerative modeling under quadratic optimal transport (OT) aims to learn deterministic maps that push mass from a simple source distribution \(p_0\) to a target distribution \(p_1\) along the Wasserstein-2 (W2) geodesics. While flow-based models and neural differential equations offer flexible transports, existing approaches typically rely on multi-step integration and yield trajectories whose curvature deviates from W2 geodesics, reducing efficiency, interpretability, and stability. We propose a geometry-aware framework that parameterizes time-dependent velocity fields as gradients of convex potentials modeled by Input Convex Neural Networks (ICNNs). This convex-potential representation guarantees transport along straight lines, exactly matching the W2 map under quadratic cost. Training uses a Flow Matching objective tailored to the convex setting, with explicit gradient computations and a dedicated inversion subproblem to recover preimages under the convex-potential flow; an optional amortization network provides favorable initializations for the inversion and accelerates optimization. The method is agnostic to the specific transport plan and can condition on arbitrary couplings between \(p_0\) and \(p_1\). Empirically, the approach yields geometry-faithful transports along W2 geodesics, enabling fast sampling with one-step or few-step updates and controlled curvature. Diagnostics on representative datasets confirm geometric fidelity and trainability, and we discuss initialization and transport-plan considerations for scalable, stable generative modeling under quadratic OT.
-
2510.0025ViewBeyond Essence: HUMN-DEF’s Seven-Axis Map of Scholarly Definitions of “the Human”Definitions of the human span biology, psychology, anthropology, law, and philosophy, resisting reduction to a single trait. This study introduces HUMN-DEF, a multiaxial framework that models seven definitional axes—Taxonomic/Evolutionary (A1), Genetic/Developmental (A2), Cognitive/Linguistic (A3), Physiological/Regulatory (A4), Sociocultural/Anthropological (A5), Legal/Normative (A6), and Phenomenological/Subjective (A7)—and represents texts as Definition Profile Vectors (DPVs). A purposive cross-disciplinary corpus (n = 31) was coded by two independent automated procedures (Krippendorff’s α = .84), analyzed with post-stratification weights (field × decade × language), and evaluated via percentile bootstraps. Results converge on Sociocultural (A5) and Cognitive/Linguistic (A3) as predominant emphases; Taxonomy/Genetics (A1/A2) anchor but are not sufficient; Legal/Normative (A6) rises under balanced representation; Phenomenology (A7) is mid-level; Physiology (A4) is specialized. Cross-field disagreement, measured with a Definitional Diversity Index (Jensen–Shannon divergence), is moderate (0.394; 95% CIs ≈ [0.345, 0.475]). We argue that “human” is best treated as a transparent, context-weighted mixture over A1–A7.
-
2510.0024ViewLECTOR: LLM-Enhanced Concept-based Test-Oriented RepetitionSpaced repetition systems are fundamental to efficient learning and memory retention, but existing algorithms often struggle with semantic interference and personalized adaptation. We present LECTOR (\textbf{L}LM-\textbf{E}nhanced \textbf{C}oncept-based \textbf{T}est-\textbf{O}riented \textbf{R}epetition), a novel adaptive scheduling algorithm specifically designed for test-oriented learning scenarios, particularly language examinations where success rate is paramount. LECTOR leverages large language models for semantic analysis while incorporating personalized learning profiles, addressing the critical challenge of semantic confusion in vocabulary learning by utilizing LLM-powered semantic similarity assessment and integrating it with established spaced repetition principles. Our comprehensive evaluation against six baseline algorithms (SSP-MMC, SM2, HLR, FSRS, ANKI, THRESHOLD) across 100 simulated learners over 100 days demonstrates significant improvements: LECTOR achieves a 90.2\% success rate compared to 88.4\% for the best baseline (SSP-MMC), representing a 2.0\% relative improvement. The algorithm shows particular strength in handling semantically similar concepts, reducing confusion-induced errors while maintaining computational efficiency. Our results establish LECTOR as a promising direction for intelligent tutoring systems and adaptive learning platforms.