All Articles
19349 articles total
YeasierAgent: Agentic Social Sandbox as a Canvas for Intent-Driven Creation of Platform-Agnostic Symbiotic Agent-Native Applications
Introduces YeasierAgent, a paradigm for creating platform-agnostic, agent-native applications using symbiotic agents and narrative worlds.
TwinBI: An Agentic Digital Twin for Efficient Augmented Interactions with Business Intelligence Dashboards
Presents TwinBI, an agentic digital-twin framework that synchronizes LLM assistance with BI dashboard states to improve analytical accuracy.
When Sample Selection Bias Precipitates Model Collapse
Research showing that sample selection bias in low-resource data silos can actually accelerate model collapse when training on synthetic data.
AI Receptivity or AI Adoption Breadth? A Tool-Specific Reanalysis of the Lower-Literacy/Higher-Usage Link
A reanalysis of AI adoption suggesting that lower AI literacy specifically predicts higher usage of non-text AI tools, rather than general AI receptivity.
MA-ProofBench: A Two-Tiered Evaluation of LLMs for Theorem Proving in Mathematical Analysis
Introduces MA-ProofBench, the first formal theorem-proving benchmark for Mathematical Analysis, revealing significant weaknesses in current LLMs.
Poker Arena: Multi-Axis Profiling of Strategic Reasoning and Memory in LLMs
Presents Poker Arena, a tournament platform to profile LLMs' strategic reasoning and memory using a multi-axis cognitive profile.
Hyperdimensional computing for structured querying on tabular data embeddings
Investigates Hyperdimensional Computing (HDC) for tabular data embeddings to provide interpretable similarity scores for structured querying.
Capability Minimization as a Safety Primitive: Risk-Aware Causal Gating for Least-Privilege LLM Agents
Introduces Risk-Aware Causal Gating (RACG) to create safer LLM agents by gating decisions based on counterfactual risk rather than confidence.
A Multi-Agent AI System for Automated High School Transcript Processing: Collaborative Document Analysis at Scale
Develops a multi-agent AI system to automate high school transcript processing with 96.7% accuracy.
Sorries Are Not the Hard Part: An Expert-Review Case Study of a Semi-Autonomous Formalization
A case study arguing that autoformalization should be judged by expert review of API design and definitions, not just whether the proof compiles.
Prove you're human by winning a claw machine
A discussion on using a claw machine as a human-verification mechanism.
David Sacks on Anthropic export control
David Sacks discusses export controls related to Anthropic's AI operations.
Orbio raises $21 million to automate hiring and onboarding for frontline workers
Orbio raises $21 million in Series A funding to automate hiring and onboarding for frontline workers.
A Deep Reinforcement Learning (DRL)-Based Transformer Method for Solving the Open Shop Scheduling Problem
Researchers propose a Transformer-based scheduling policy to solve the Open Shop Scheduling Problem, showing strong generalization from small to large instances.
UP-NRPA: User Portrait based Nested Rollout Policy Adaptation for Planning with Large Language Models in Goal-oriented Dialogue Systems
UP-NRPA is introduced as an online framework for goal-oriented dialogue systems that adapts to user portraits without offline RL.
History of the Muddy Children Puzzle
A historical survey of the Muddy Children Puzzle and its influence on epistemic logic, including a new self-referential hats puzzle.
Orchestra-o1: Omnimodal Agent Orchestration
Orchestra-o1 is presented as an omnimodal agent orchestration framework using DA-GRPO for efficient training of multi-agent collaboration.
Hybrid Open-Ended Tri-Evolution Makes Better Deep Researcher
The HOTE framework utilizes hybrid-mode reinforcement learning to evolve proposer, solver, and judge modules for deep research agents.
WorkBench Revisited: Workplace Agents Two Years On
A longitudinal study of the WorkBench benchmark shows significant improvements in agent capability and safety between 2024 and 2026.
Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP
A comparison of Diff-in-Means and Iterative Nullspace Projection for steering refusal in open-weight chat models.