[{"id":"cms76jv1tatytkh0c6g7t0o5i","channel":"knowledge","topic":"arxiv-ai","title":"AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents","summary":"Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools. This dependence allows defenders to inject deceptive observations that can mislead the agent's decision-making process. However, existing defe","payload":{"url":"https://arxiv.org/abs/2607.26998v1","title":"AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.26998","version":"1","abstract":"Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools. This dependence allows defenders to inject deceptive observations that can mislead the agent's decision-making process. However, existing defenses rely heavily on static, isolated artifacts planted in the environment prior to an attack. Advanced agents can progressively recognize and bypass these artifacts, ultimately refocusing their exploitation attempts on the real target. To address this issue, we introduce AgentSnare, a trajectory-adaptive deception system that dynamically unfolds a decoy environment to continually steer the penetration agent away from the real target. Specifically, AgentSnare employs an artifact-construction policy model that constructs candidate artifacts conditioned on the agent's interaction history and decoy state. AgentSnare then validates these candidates and incrementally incorporates valid artifacts into a factually consistent decoy environment, thereby delaying the attack by absorbing its tool calls, diverting its post-entry trajectory within the decoy, and defusing it by inducing completion rep","arxiv_id":"2607.26998","categories":["cs.CR","cs.CL","cs.LG"],"published_at":"2026-07-29T14:56:31.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:23.153Z"},{"id":"cms76jujdatyrkh0cdf1oj28j","channel":"knowledge","topic":"arxiv-ai","title":"Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making","summary":"Regional bias in large language models (LLMs) may shape both perceptions of regional groups and decisions about individuals from different regions. Yet existing studies often examine these manifestations separately, leaving their structure and consequences unclear. We introduce Stereotypes-to-Decisi","payload":{"url":"https://arxiv.org/abs/2607.27022v1","title":"Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27022","version":"1","abstract":"Regional bias in large language models (LLMs) may shape both perceptions of regional groups and decisions about individuals from different regions. Yet existing studies often examine these manifestations separately, leaving their structure and consequences unclear. We introduce Stereotypes-to-Decisions (S2D), a systematic framework evaluating regional bias from abstract stereotypes to concrete social decisions. Covering all 34 provincial-level administrative regions of China, S2D evaluates six LLMs using stereotype ratings of Warmth (perceived friendliness and trustworthiness) and Competence (perceived capability and intelligence), along with paired-choice tasks across Education, Occupation, and Social Interaction. Results reveal substantial regional differences in regional scores, with considerable agreement across models, especially for Competence and Occupation decisions. Furthermore, these patterns are associated with regional economic and digital development indicators and display mixed human-like stereotypes, with some regions rated highly on one dimension but poorly on the other. They also remain largely stable across Chinese and English prompts. Overall, our findings show t","arxiv_id":"2607.27022","categories":["cs.CL"],"published_at":"2026-07-29T15:20:36.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:22.489Z"},{"id":"cms76ju0ratypkh0c7bhw2dy6","channel":"knowledge","topic":"arxiv-ai","title":"BayesAME: Bayesian Active Model Evaluation","summary":"Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items, known as a coreset. Current literature mostly requires the practitioner ","payload":{"url":"https://arxiv.org/abs/2607.27023v1","title":"BayesAME: Bayesian Active Model Evaluation","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27023","version":"1","abstract":"Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items, known as a coreset. Current literature mostly requires the practitioner to input a coreset size. However, when reliable performance estimation takes priority over efficiency, an evaluation method should also be capable of automatically determining a coreset size that reflects this priority. We introduce BayesAME, a sequential Bayesian framework specifically targeting automatic determination of the coreset size. BayesAME models performance as a random variable by defining a latent ability for each group of items sharing the same historical model performances, with a joint prior distribution encoding the belief that the target model behaves similarly to these historical models. The posterior distribution over these abilities is used to derive performance estimators, quantify performance uncertainty, and select items to add to the coreset via an information-gain criterion. The coreset is iteratively augmented until the performance estimate fluctuation and the p","arxiv_id":"2607.27023","categories":["cs.LG","cs.AI","stat.ML"],"published_at":"2026-07-29T15:20:39.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:21.820Z"},{"id":"cms76jthjatynkh0clu3dm7ni","channel":"knowledge","topic":"arxiv-ai","title":"CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation","summary":"Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. However, differences in architectural indu","payload":{"url":"https://arxiv.org/abs/2607.27054v1","title":"CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27054","version":"1","abstract":"Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model compression. The emergence of diverse model architectures has extended KD from homogeneous to heterogeneous settings. However, differences in architectural inductive biases between the teacher and student models often result in substantial representation discrepancies, limiting the effectiveness of direct knowledge transfer. Recently, redundancy suppression has offered a new perspective on heterogeneous KD by preserving cross-architecture invariance and reducing feature redundancy through decorrelation of teacher-student feature correlations. Nevertheless, this formulation may weaken useful structural information through uniform decorrelation, while a fixed coefficient may make the effective contribution of redundancy suppression sensitive to teacher-student pairs and training stages. To address these problems, Correlation Calibration-based Redundancy Suppression (CoCaRS) is proposed to better retain structural information while suppressing redundancy and reduce sensitivity to coefficient settings across teacher-student pairs and training stage","arxiv_id":"2607.27054","categories":["cs.LG","cs.AI"],"published_at":"2026-07-29T15:47:11.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:21.127Z"},{"id":"cms76jsyuatylkh0cgipuubr4","channel":"knowledge","topic":"arxiv-ai","title":"Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data","summary":"Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory bench","payload":{"url":"https://arxiv.org/abs/2607.27056v1","title":"Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27056","version":"1","abstract":"Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversational histories, failing to provide an effective assessment of deeper user understanding. In this work, we propose Setoka, a benchmark for evaluating memory-augmented personalized agents with hierarchical user understanding from heterogeneous data. Grounded in theories from cognitive and personality psychology, Setoka defines four levels of user understanding, i.e., semantic memory, episodic memory, behavior pattern, and personality trait. Moreover, to enable realistic yet privacy-preserving evaluation, we design a psychometrics-based pipeline that synthesizes diverse, coherent heterogeneous user data and queries at scale. Finally, we leverage Setoka to evaluate 3 language models combined with 5 memory systems for 10 synthetic users. Our comprehensive evaluation reveals that while existing sy","arxiv_id":"2607.27056","categories":["cs.AI","cs.CL"],"published_at":"2026-07-29T15:47:40.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:20.454Z"},{"id":"cms76jsg4atyjkh0coiii9nus","channel":"knowledge","topic":"arxiv-ai","title":"ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection","summary":"While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarcity of annotated defect data make this task challenging. This paper presents a procedural rendering pipeline that generates large-scale annotated synthetic train","payload":{"url":"https://arxiv.org/abs/2607.27065v1","title":"ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27065","version":"1","abstract":"While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarcity of annotated defect data make this task challenging. This paper presents a procedural rendering pipeline that generates large-scale annotated synthetic training data using BlenderProc, with configurable material appearance, camera modes, and domain randomization, producing automatic COCO-format annotations. To show the potential of our approach, we evaluate four training strategies, namely synthetic-only, real-only, mixed, and fine-tuning from synthetic weights, across two objects with different material properties and three lightweight edge-deployable detectors, YOLOX, YOLO26, and LW-DETR. Our evaluation show that fine-tuning from synthetic weights consistently outperforms real-only training, and that mixed training effectively recovers performance under scarce real-data conditions, with findings validated across both convolutional and transformer-based architectures. The proposed approach enables scalable defect detection without the burden of large real annotated datasets, making it practical for on-device industrial inspection. The pipel","arxiv_id":"2607.27065","categories":["cs.CV","cs.AI"],"published_at":"2026-07-29T15:54:01.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:19.780Z"},{"id":"cms76jrx7atyhkh0cpmgu6ftj","channel":"knowledge","topic":"arxiv-ai","title":"SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence","summary":"Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully support the manuscript's claims, and communicate evidence with a clear visual hierarchy. However, if we apply traditional image assessment methods to","payload":{"url":"https://arxiv.org/abs/2607.27066v1","title":"SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27066","version":"1","abstract":"Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully support the manuscript's claims, and communicate evidence with a clear visual hierarchy. However, if we apply traditional image assessment methods to scientific figure quality assessment, limitations emerge: classic IQA models capture perceptual quality or aesthetics but cannot judge whether a figure serves the paper's scientific argument; CLIP-based methods assess generic image-text correspondence, yet lack understanding of manuscript context; and zero-shot LLM/VLM judges, when repurposed for figure scoring, often yield overly concentrated scores with limited fusion of visual and textual evidence. We introduce an annotated dataset of 3,857 scientific figures from peer-reviewed conference papers, each rated along four peer-review-oriented dimensions: Clarity, Relevance, Informativeness, and Structure. We propose SciFigAlign, a fine-tuned multimodal scorer that grounds figure quality assessment in manuscript evidence. Given a figure crop, caption, citing paragraphs, and light paper context, SciFigAlign fine-tunes CLIP and SciBERT end-","arxiv_id":"2607.27066","categories":["cs.CV","cs.AI"],"published_at":"2026-07-29T15:54:14.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:19.099Z"},{"id":"cms76jre4atyfkh0cr2wtxhkl","channel":"knowledge","topic":"arxiv-ai","title":"Visual Credit Audit for Multimodal Spatial Reasoning","summary":"Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decision more support than t","payload":{"url":"https://arxiv.org/abs/2607.27069v1","title":"Visual Credit Audit for Multimodal Spatial Reasoning","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27069","version":"1","abstract":"Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decision more support than text-only and blank controls, and whether the model responds to relation-specific visual evidence. The first audit is training- and label-free and does not require an answer flip. Applying labels yields dependence-credited correctness (D-CC); on correct items, it equals same-control gold-aligned positive gain, while prediction alignment extends the audit to errors. Across four open MLLMs and two spatial benchmarks, 12.73-26.25% of decisions are correct yet uncredited. Matched same-split image permutation reduces D-CC by 21.25-47.80 points, with every paired 95% interval above zero. Fixed-pixel relation contrasts and a 3x3 evidence-source factorial show why null controls cannot identify relation response. Among controlled correct-but-uncredited agreement decisions, response to relation reversal spans 81.57-100.00%, while 32.11% pooled change answer. Independently audited outcomes on 108 ge","arxiv_id":"2607.27069","categories":["cs.CV","cs.AI"],"published_at":"2026-07-29T15:55:31.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:18.412Z"},{"id":"cms76jqv4atydkh0cfovmczdp","channel":"knowledge","topic":"arxiv-ai","title":"Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise","summary":"We study online convex optimization (OCO) in non-stationary environments under heavy-tailed noise, where the stochastic gradient oracle admits only a finite $p$-th central moment for some $p \\in (1, 2]$. While static regret is well-understood, achieving universal dynamic regret in a parameter-free m","payload":{"url":"https://arxiv.org/abs/2607.27073v1","title":"Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27073","version":"1","abstract":"We study online convex optimization (OCO) in non-stationary environments under heavy-tailed noise, where the stochastic gradient oracle admits only a finite $p$-th central moment for some $p \\in (1, 2]$. While static regret is well-understood, achieving universal dynamic regret in a parameter-free manner remains an open challenge. We resolve this by proposing \\textbf{HT-PAder}, a parameter-free algorithm combining restarted AdaGrad experts over a geometric pool of block lengths with a pathwise meta-algorithm, \\textbf{AdaGrad-Hedge}, which requires no moment conditions on meta-losses. For a domain of diameter $D$, Lipschitz constant $G$, noise level $σ$, and comparator path length $P_T$, HT-PAder achieves an expected universal dynamic regret of \\[ \\widetilde O\\left( GD\\sqrt{T(1+P_T/D)} + σD T^{1/p}(1+P_T/D)^{(p-1)/p} \\right). \\] The algorithm does not require prior knowledge of any of these problem parameters. Even in the special case of finite variance ($p=2$), HT-PAder provides the first parameter-free minimax universal dynamic regret guarantee. We also prove a matching lower bound, establishing the optimality of the path-length exponent.","arxiv_id":"2607.27073","categories":["cs.LG","cs.AI","math.OC"],"published_at":"2026-07-29T15:58:18.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:17.728Z"},{"id":"cms76jqckatybkh0cju95cut3","channel":"knowledge","topic":"arxiv-ai","title":"MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair","summary":"Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A malicious instruction crafted by an attacker may be stored in long-term memory, recalled much later, and quietly shape a real action. Recent benchmarks increasingly ","payload":{"url":"https://arxiv.org/abs/2607.27080v1","title":"MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27080","version":"1","abstract":"Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A malicious instruction crafted by an attacker may be stored in long-term memory, recalled much later, and quietly shape a real action. Recent benchmarks increasingly examine agent memory security, yet few trace the same malicious semantics across persistence, downstream consequences, and selective repair under diverse memory-backend comparisons. To address this gap, we introduce MemSecBench, a task-grounded benchmark for the lifecycle security of agent memory systems. It contains 310 cases drawn from 48 realistic contexts across code and science, daily life, and office work. Each case follows a controlled Write--Execute--Forget protocol in an isolated runtime under an exact agent configuration, defined by an agent harness, a memory backend, and an LLM backend. Evidence-based adjudication combines a deterministic write check, checkpoint-specific judge-model evaluations, and programmatic gates across seven lifecycle checkpoints. The experimental design spans a 24-configuration matrix of two agent harnesses, four memory backends, and three LLM backends.","arxiv_id":"2607.27080","categories":["cs.CR","cs.AI"],"published_at":"2026-07-29T16:06:54.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:17.059Z"},{"id":"cms76jpu4aty9kh0cabwtbxjf","channel":"knowledge","topic":"arxiv-ai","title":"On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment","summary":"Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand. Existing ","payload":{"url":"https://arxiv.org/abs/2607.27081v1","title":"On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27081","version":"1","abstract":"Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand. Existing safety-realignment defenses often fail in practice due to three key limitations: they frequently cause catastrophic forgetting of specialized skills; their effectiveness collapses when the defender cannot observe the attacker's prompt template; and successfully realigned models remain susceptible to re-jailbreaking via simple system prompt switches. To address these challenges, we propose Routing-based On-Policy Distillation (ROPD), a novel realignment framework that models the divergence between aligned and compromised output probability distributions rather than fitting specific prompt templates. We conduct extensive experiments comparing ROPD against four state-of-the-art baselines across three datasets and three base models with varying alignment strengths. Our results demonstrate that when baseline defenses face template mismatches, often accompanied by severe degradation in downstr","arxiv_id":"2607.27081","categories":["cs.AI","cs.CL","cs.CR","cs.LG"],"published_at":"2026-07-29T16:07:19.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:16.396Z"},{"id":"cms76jpbiaty5kh0c0wujm9f8","channel":"knowledge","topic":"arxiv-ai","title":"Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents","summary":"As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, while too many adds cost, context load, and privacy exposure. Routers","payload":{"url":"https://arxiv.org/abs/2607.27083v1","title":"Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27083","version":"1","abstract":"As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fundamental tool-selection challenge: acquiring too few tools leaves the task under-informed, while too many adds cost, context load, and privacy exposure. Routers and retrievers can rank candidate tools by relevance, but a ranking alone does not determine how many are worth selecting. Existing approaches leave acquisition under heterogeneous costs unaddressed. We formulate this decision as cost-aware marginal decision-focused stopping (CAM-DF) over ranked tool prefixes, with CAM-DF-lite as a compact interpretable variant. We train directly on the offline gap between stopping now and the best continuation: its sign labels the decision, its magnitude weights each error by the payoff at stake. We prove this objective is Bayes-aligned with the stopping target and that score-only rules are suboptimal under heterogeneous costs. We evaluate on 1,343 tasks across five tool-use domains. On $τ$-bench Retail, CAM-DF attains the highest payoff among deployable methods, with gains over a predict-then-threshold baseline across all five ranking sources and two ","arxiv_id":"2607.27083","categories":["cs.LG","cs.AI"],"published_at":"2026-07-29T16:07:37.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:15.726Z"},{"id":"cms76joryaty1kh0cfkuj59np","channel":"knowledge","topic":"arxiv-ai","title":"SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context","summary":"Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated cont","payload":{"url":"https://arxiv.org/abs/2607.27084v1","title":"SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27084","version":"1","abstract":"Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers. The few existing studies on scholarly charts remain confined to visual-surface comparisons, failing to verify caption alignment, citation relevance, or visual misleadingness. To address this, we propose SciFigQual-Bench, a full-text contextual benchmark that evaluates scientific images across five dimensions (clarity, layout, caption fit, context relevance, and misleading risk). The data covers top computer-science conferences from 2020 to 2025; 6,308 images were independently scored by multiple domain experts in five dimensions and aggregated into gold-standard annotations. Unlike previous scientific figure benchmarks, our dataset binds each image to its caption, citing sentence, and manuscript context. To enable automated evaluation on this benchmark, we designed a staged cross-modal evaluation framework SFQ-Agent to achieve a","arxiv_id":"2607.27084","categories":["cs.CV","cs.AI"],"published_at":"2026-07-29T16:07:44.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:15.022Z"},{"id":"cms76jo9matxzkh0cquhapauz","channel":"knowledge","topic":"arxiv-ai","title":"MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning","summary":"With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information c","payload":{"url":"https://arxiv.org/abs/2607.27109v1","title":"MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27109","version":"1","abstract":"With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \\textbf{M}assive \\textbf{M}ulti-dimensional benchmark for \\textbf{A}udio \\textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.","arxiv_id":"2607.27109","categories":["cs.SD","cs.AI"],"published_at":"2026-07-29T16:38:08.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:14.362Z"},{"id":"cms76jnr2atxxkh0cfwxvp3ms","channel":"knowledge","topic":"arxiv-ai","title":"AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching","summary":"Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems identify only one type of semantic correspondence and cannot simultaneously discover equivalence and subsumption mappings. In this paper, we introduce Hybrid Onto","payload":{"url":"https://arxiv.org/abs/2607.27130v1","title":"AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27130","version":"1","abstract":"Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems identify only one type of semantic correspondence and cannot simultaneously discover equivalence and subsumption mappings. In this paper, we introduce Hybrid Ontology Matching (HOM), a new OM task that unifies equivalence and subsumption discovery, and accordingly propose a Large Language Model (LLM)-based multi-agent OM framework AgentMap that is implemented by a series of interdependent semantic decisions. Given a concept in the source ontology, AgentMap integrates semantic retrieval, hierarchical search, and collaborative multi-agent LLM reasoning to progressively explore the target ontology, identifying either the equivalent concept, if one exists, or the most fine-grained subsumer. We further extend four OM datasets for a HOM benchmark and evaluate AgentMap under hybrid, equivalence-only, and subsumption-only settings. Experimental results show that AgentMap achieves promising performance on the hybrid setting, and at the same time outperforms equivalence matching and subsumption matching baselines on the equivalence-only and subsumption-onl","arxiv_id":"2607.27130","categories":["cs.AI"],"published_at":"2026-07-29T16:58:10.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:13.694Z"},{"id":"cms76jn8matxvkh0c9tlblyhk","channel":"knowledge","topic":"arxiv-ai","title":"Linguistic Monoculture in LLM-Assisted Language Use","summary":"Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduce population-level va","payload":{"url":"https://arxiv.org/abs/2607.27134v1","title":"Linguistic Monoculture in LLM-Assisted Language Use","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27134","version":"1","abstract":"Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduce population-level variation in linguistic form, a phenomenon we refer to as linguistic monoculture. We develop a mathematical framework in which authors and LLMs are represented as distributions over linguistic features and coevolve through repeated interaction. We analyze three interaction mechanisms: a shared model with a fixed linguistic distribution, a shared model recursively updated from author outputs, and personalized models updated through author-specific and population-level feedback. We characterize the resulting equilibria and convergence rates, showing that, shared models can drive authors toward a common norm, recursive feedback relocates the shared norm without altering pairwise spread under common conformity, and personalization can preserve a family of distinct author-model equilibria with nonzero linguistic diversity. We then endogenize conformity as a strategic choice trading off private ","arxiv_id":"2607.27134","categories":["cs.AI","cs.CL","cs.GT"],"published_at":"2026-07-29T17:04:36.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:13.030Z"},{"id":"cms76jmq5atxrkh0cnfksyfkz","channel":"knowledge","topic":"arxiv-ai","title":"DLAM: Distributional Latent Actions with Temporal Constraints","summary":"Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure ","payload":{"url":"https://arxiv.org/abs/2607.27138v1","title":"DLAM: Distributional Latent Actions with Temporal Constraints","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27138","version":"1","abstract":"Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure required for joint generation with robot actions. Existing structured methods add temporal constraints but retain deterministic transition points, so residual errors in locally inferred transitions may propagate and compound under recursive composition. We introduce DLAM, a distributional latent-action model that represents each transition as a diagonal Gaussian. Reconstruction conditioned on the reference frame grounds the mean in observed visual change, while normalized composition and reversal over equal-gap triplets constrain both the mean and dimension-wise variance. Variance composition uses a lightweight shared-correlation coefficient to account for dependence between adjacent transitions that share an intermediate frame, whereas reversal negates the mean and preserves the variance. For downstream policy learning, we freeze the encoder and train a flow-matching policy to jointly g","arxiv_id":"2607.27138","categories":["cs.RO","cs.AI","cs.CV"],"published_at":"2026-07-29T17:09:48.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:12.365Z"},{"id":"cms76jm7ratxnkh0ctb5qv6d7","channel":"knowledge","topic":"arxiv-ai","title":"Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark","summary":"High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal prediction (CP) provides valid overall coverage guarantees; however, we ","payload":{"url":"https://arxiv.org/abs/2607.27143v1","title":"Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27143","version":"1","abstract":"High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal prediction (CP) provides valid overall coverage guarantees; however, we show that it severely under-covers rare, costly minority classes, with minority-class coverage dropping to as low as 0.5% on certain datasets. To characterize and address this limitation, we conduct a comprehensive benchmark comparing marginal CP, class-conditional (Mondrian) CP, and cost-controlled abstention mechanisms across 15 real-world imbalanced tabular datasets, 7 classification models, 3 probability calibration techniques, and 10 random seeds, resulting in 3,150 experimental runs. Our results show that Mondrian CP restores valid minority-class coverage, achieving an average minority-coverage improvement of 61.7 percentage points over marginal CP (p < 1e-80). Furthermore, combining Mondrian CP with cost-controlled abstention significantly reduces expected decision cost compared with standard decision boundaries, confidence-based rejectors, and risk-controlled rejectors under real","arxiv_id":"2607.27143","categories":["cs.LG","cs.AI"],"published_at":"2026-07-29T17:15:33.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:11.704Z"},{"id":"cms76jlpjatxjkh0cr8svqron","channel":"knowledge","topic":"arxiv-ai","title":"MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis","summary":"Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolv","payload":{"url":"https://arxiv.org/abs/2607.27146v1","title":"MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27146","version":"1","abstract":"Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolve fewer than 1% of tasks. One obstacle is the lack of scalable training environments for this from-scratch setting, spanning the whole software engineering life cycle, as existing environment-construction frameworks focus only on a single phase in software development. To address this gap, we introduce MindForge, an automated pipeline that converts open-source command-line programs into source-free environments that expose only a compiled reference executable and its documentation. Using MindForge, we construct training environments from repositories disjoint from those in ProgramBench, and curate a high-quality data recipe consisting of program synthesis trajectories using GLM-5.2 as the teacher agent. Fine-tuning Qwen3.6-27B on these trajectories increases its ProgramBench average test pass rate from 37.98% to 49.51%, achieving performance comparable to substantially larger frontier mo","arxiv_id":"2607.27146","categories":["cs.SE","cs.CL","cs.LG"],"published_at":"2026-07-29T17:23:02.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:11.047Z"},{"id":"cms76jl72atxfkh0c46ueoi64","channel":"knowledge","topic":"arxiv-ai","title":"Anatomy Contextualized Adaption of CT Foundation Models","summary":"CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning anatomy-level visual fea","payload":{"url":"https://arxiv.org/abs/2607.27154v1","title":"Anatomy Contextualized Adaption of CT Foundation Models","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27154","version":"1","abstract":"CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning anatomy-level visual features with anatomy-specific text, but in doing so discards the global context that whole-volume models provide. Furthermore, existing fine-grained approaches train from scratch, making them computationally expensive. We introduce Anatomy Contextualized Adaptation (ACA), a lightweight framework that adapts frozen CT foundation model representations for anatomy-level vision-language alignment while enhancing global contextualization. ACA uses TotalSegmentator to decompose CT volumes into anatomy-level embeddings, which are refined via a transformer that captures cross-anatomy relationships, and aligned to both per-anatomy and scan-level text extracted from radiology reports. Evaluated on Merlin and CT-RATE, ACA consistently outperforms both the frozen foundation model baselines and existing fine-grained methods in zero-shot finding classification, while requiring less than one hour of trai","arxiv_id":"2607.27154","categories":["cs.CV","cs.AI"],"published_at":"2026-07-29T17:32:57.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:10.382Z"},{"id":"cms76jkohatxdkh0cqm6nojra","channel":"knowledge","topic":"arxiv-ai","title":"OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding","summary":"Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating L","payload":{"url":"https://arxiv.org/abs/2607.27155v1","title":"OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27155","version":"1","abstract":"Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation. To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality. The code and dataset are fully open-sourced, a","arxiv_id":"2607.27155","categories":["cs.AI","cs.CL","cs.HC"],"published_at":"2026-07-29T17:33:47.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:09.713Z"},{"id":"cms76jk67atxbkh0cx1q38ca2","channel":"knowledge","topic":"arxiv-ai","title":"SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch","summary":"LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execute-only binary as a ","payload":{"url":"https://arxiv.org/abs/2607.27167v1","title":"SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27167","version":"1","abstract":"LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execute-only binary as a behavioral oracle, even frontier models solve fewer than 1% of instances. Existing frameworks conflate documentation reading, behavioral exploration, and code synthesis into a single pass, causing agents to probe insufficiently, lose behavioral intent as context drifts, and propagate early misinterpretations into the final implementation. Inspired by classical requirements engineering, we argue that behavioral specification elicitation should be a first-class phase that precedes implementation. We present SpecFirst, a two-stage framework that forces the specification elicitation before code synthesis. A dedicated spec agent first probes the binary and combines observations with documentation into a structured specification. Next, a code synthesis agent then uses this specification to drive implementation. This decomposition resolves documentation ambiguities before coding begins and prov","arxiv_id":"2607.27167","categories":["cs.SE","cs.CL"],"published_at":"2026-07-29T17:42:47.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:09.056Z"},{"id":"cms76jjntatx9kh0cibw1i1eh","channel":"knowledge","topic":"arxiv-ai","title":"Improving Item Discoverability in e-Commerce Search via Related Intent Generation","summary":"Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-commerce marketplaces and particularly grocery, this paradigm is limiting, as user satisfaction and commercial outcomes depend heavily on the discoverability of subs","payload":{"url":"https://arxiv.org/abs/2607.27172v1","title":"Improving Item Discoverability in e-Commerce Search via Related Intent Generation","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27172","version":"1","abstract":"Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-commerce marketplaces and particularly grocery, this paradigm is limiting, as user satisfaction and commercial outcomes depend heavily on the discoverability of substitute, complementary, and thematically related items. In this paper, we present a scalable system for discovery-augmented search that leverages intent-conditioned recall expansion. Our approach generates implicit user intents to expand candidate recall while maintaining relevance. The system addresses the cost-quality tradeoff of generative retrieval through a two-stage hybrid architecture. First, we leverage closed-weight large language models (LLMs) to maximize discoverability for head queries. To extend these benefits to tail queries, we then introduce a finetuned small language model (SLM), trained via LoRA adapters and teacher-student distillation. We evaluate the system using a rigorous dual framework: (a) LLM-as-a-judge metrics validated against human preferences for semantic quality, and (b) end-to-end session-level purchase analysis. Results demonstrate that our approach improv","arxiv_id":"2607.27172","categories":["cs.IR","cs.AI"],"published_at":"2026-07-29T17:46:35.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:08.393Z"},{"id":"cms76jj53atx7kh0cvbmhg55q","channel":"knowledge","topic":"arxiv-ai","title":"Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork","summary":"Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed task and that the partner's capabilities, their ability to successfully execute the desired action, a","payload":{"url":"https://arxiv.org/abs/2607.27177v1","title":"Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27177","version":"1","abstract":"Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed task and that the partner's capabilities, their ability to successfully execute the desired action, are already known. In reality, a partner's true capabilities are often hidden, and human collaborators may act sub-optimally on tasks with multiple valid strategies. To address these limitations, we extend ad-hoc teamwork into a multi-task setting by re-framing it as a problem of joint planning with decentralised execution under hidden partner capabilities. We introduce CE-CM (Capability Estimation via Contextual Models), an approximate Bayesian method that infers task-invariant capability vectors. By using simulation-based sampling, the agent estimates capabilities and induces a contextual Multi-agent Markov Decision Processes for planning. This approach requires no population pre-training and refines its beliefs online from just a few tasks. To account for human unpredictability, we propose CE-CM-Div, an extension that evaluates capability hypotheses against diverse planner rollouts rat","arxiv_id":"2607.27177","categories":["cs.AI","cs.HC","cs.MA"],"published_at":"2026-07-29T17:50:39.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:07.719Z"},{"id":"cms76jim3atx5kh0co32yyi92","channel":"knowledge","topic":"arxiv-ai","title":"DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search","summary":"State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first reconstruct and cura","payload":{"url":"https://arxiv.org/abs/2607.27178v1","title":"DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search","source":"arxiv","pdf_url":"https://arxiv.org/pdf/2607.27178","version":"1","abstract":"State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first reconstruct and curate 665M English contrastive pre-training pairs from 1.4B pairs across 34 public sources and build 1.88M supervised fine-tuning pairs with mined hard negatives. Training yields two 149M-parameter models: DenseOn, a single-vector dense model, and LateOn, a ColBERT-style late-interaction model. They achieve 56.20 and 57.22 average nDCG@10 on BEIR, respectively, setting new state-of-the-art results for this size class. We then translate the validated English data into eight languages, yielding 2.8B pairs with cross-lingual samples, and train mDenseOn and mLateOn, two 307M-parameter models built on mmBERT-base. Despite sharing their backbone, data, and objectives, their representations behave differently: the dense model is strong on English and translated languages but degrades outside translate-train support, whereas the late-interaction model generalizes better to unseen languages and scri","arxiv_id":"2607.27178","categories":["cs.CL","cs.IR"],"published_at":"2026-07-29T17:50:51.000Z"},"public_metadata":null,"published_at":"2026-07-30T07:16:07.035Z"}]