All Articles
19339 articles total
Why does paper fold so well?
A discussion on the physical properties of paper and why it folds efficiently.
Adversarial Concept Search: Predicting Compositional Errors From Feature Geometry
Researchers propose using representational geometry in LLMs to predict compositional failures and identify high-risk scenarios without evaluating specific inputs.
Minim: Privacy-Aware Minimal View for Agents via Trusted Local Sanitization
MINIM is a trusted local broker that minimizes UI state observations on the client side to prevent sensitive data leakage when using LLM agents.
Formalizing Numerical Analysis: An Agent Pipeline and Quality Audit Beyond Kernel Acceptance
A study on using coding agents to formalize numerical analysis textbooks in Lean 4, introducing a framework to evaluate formalization quality beyond simple kernel acceptance.
Applicability Condition Extraction for Therapeutic Drug-Disease Relations
Introduction of a dataset and method to extract applicability conditions for therapeutic drug-disease relations from biomedical literature.
FactoryLLM: A Safe and Open-Source AI Playground for Evaluating LLMs in Smart Factories
FactoryLLM is an open-source AI playground for evaluating RAG models in smart factory settings while maintaining data privacy via local LLM execution.
VeriGeo: Controllable Geometry Question Generation with Numerical and Analytical Verification
VeriGeo is a controllable geometry question generation framework that uses executable reasoning traces and verification to produce high-quality synthetic data for AI education.
When Should Agent Trust Be Conditional? Characterizing and Attacking Skill-Conditional Reputation in Agent Swarms
An analysis of skill-conditional reputation in agent swarms, demonstrating how specialized trust scores can be more effective but are vulnerable to specific hijacking attacks.
Closing the Reflection Gap: A Free Calibration Bonus for Agentic RL
RefGRPO is a reinforcement learning fix that uses a calibration bonus to close the reflection gap in LLM agents, improving their ability to self-assess performance based on environment feedback.
SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing
SkillAudit is a framework for evolving agent skills without ground-truth feedback, utilizing paired trajectory auditing and contrastive evaluation to improve specialized workflows.
YeasierAgent: Agentic Social Sandbox as a Canvas for Intent-Driven Creation of Platform-Agnostic Symbiotic Agent-Native Applications
Introduces YeasierAgent, a paradigm for creating platform-agnostic, agent-native applications using symbiotic agents and narrative worlds.
TwinBI: An Agentic Digital Twin for Efficient Augmented Interactions with Business Intelligence Dashboards
Presents TwinBI, an agentic digital-twin framework that synchronizes LLM assistance with BI dashboard states to improve analytical accuracy.
When Sample Selection Bias Precipitates Model Collapse
Research showing that sample selection bias in low-resource data silos can actually accelerate model collapse when training on synthetic data.
AI Receptivity or AI Adoption Breadth? A Tool-Specific Reanalysis of the Lower-Literacy/Higher-Usage Link
A reanalysis of AI adoption suggesting that lower AI literacy specifically predicts higher usage of non-text AI tools, rather than general AI receptivity.
MA-ProofBench: A Two-Tiered Evaluation of LLMs for Theorem Proving in Mathematical Analysis
Introduces MA-ProofBench, the first formal theorem-proving benchmark for Mathematical Analysis, revealing significant weaknesses in current LLMs.
Poker Arena: Multi-Axis Profiling of Strategic Reasoning and Memory in LLMs
Presents Poker Arena, a tournament platform to profile LLMs' strategic reasoning and memory using a multi-axis cognitive profile.
Hyperdimensional computing for structured querying on tabular data embeddings
Investigates Hyperdimensional Computing (HDC) for tabular data embeddings to provide interpretable similarity scores for structured querying.
Capability Minimization as a Safety Primitive: Risk-Aware Causal Gating for Least-Privilege LLM Agents
Introduces Risk-Aware Causal Gating (RACG) to create safer LLM agents by gating decisions based on counterfactual risk rather than confidence.
A Multi-Agent AI System for Automated High School Transcript Processing: Collaborative Document Analysis at Scale
Develops a multi-agent AI system to automate high school transcript processing with 96.7% accuracy.
Sorries Are Not the Hard Part: An Expert-Review Case Study of a Semi-Autonomous Formalization
A case study arguing that autoformalization should be judged by expert review of API design and definitions, not just whether the proof compiles.