AI/ML arXiv cs.AI

YeasierAgent: Agentic Social Sandbox as a Canvas for Intent-Driven Creation of Platform-Agnostic Symbiotic Agent-Native Applications

Introduces YeasierAgent, a paradigm for creating platform-agnostic, agent-native applications using symbiotic agents and narrative worlds.

AI/ML arXiv cs.AI

TwinBI: An Agentic Digital Twin for Efficient Augmented Interactions with Business Intelligence Dashboards

Presents TwinBI, an agentic digital-twin framework that synchronizes LLM assistance with BI dashboard states to improve analytical accuracy.

AI/ML arXiv cs.AI

When Sample Selection Bias Precipitates Model Collapse

Research showing that sample selection bias in low-resource data silos can actually accelerate model collapse when training on synthetic data.

AI/ML arXiv cs.AI

AI Receptivity or AI Adoption Breadth? A Tool-Specific Reanalysis of the Lower-Literacy/Higher-Usage Link

A reanalysis of AI adoption suggesting that lower AI literacy specifically predicts higher usage of non-text AI tools, rather than general AI receptivity.

AI/ML arXiv cs.AI

MA-ProofBench: A Two-Tiered Evaluation of LLMs for Theorem Proving in Mathematical Analysis

Introduces MA-ProofBench, the first formal theorem-proving benchmark for Mathematical Analysis, revealing significant weaknesses in current LLMs.

AI/ML arXiv cs.AI

Poker Arena: Multi-Axis Profiling of Strategic Reasoning and Memory in LLMs

Presents Poker Arena, a tournament platform to profile LLMs' strategic reasoning and memory using a multi-axis cognitive profile.

AI/ML arXiv cs.AI

Hyperdimensional computing for structured querying on tabular data embeddings

Investigates Hyperdimensional Computing (HDC) for tabular data embeddings to provide interpretable similarity scores for structured querying.

AI/ML arXiv cs.AI

Capability Minimization as a Safety Primitive: Risk-Aware Causal Gating for Least-Privilege LLM Agents

Introduces Risk-Aware Causal Gating (RACG) to create safer LLM agents by gating decisions based on counterfactual risk rather than confidence.

AI/ML arXiv cs.AI

A Multi-Agent AI System for Automated High School Transcript Processing: Collaborative Document Analysis at Scale

Develops a multi-agent AI system to automate high school transcript processing with 96.7% accuracy.

AI/ML arXiv cs.AI

Sorries Are Not the Hard Part: An Expert-Review Case Study of a Semi-Autonomous Formalization

A case study arguing that autoformalization should be judged by expert review of API design and definitions, not just whether the proof compiles.

Other Hacker News

Prove you're human by winning a claw machine

A discussion on using a claw machine as a human-verification mechanism.

Tech Business/VC Hacker News

David Sacks on Anthropic export control

David Sacks discusses export controls related to Anthropic's AI operations.

Tech Business/VC TechCrunch

Orbio raises $21 million to automate hiring and onboarding for frontline workers

Orbio raises $21 million in Series A funding to automate hiring and onboarding for frontline workers.

AI/ML arXiv cs.AI

A Deep Reinforcement Learning (DRL)-Based Transformer Method for Solving the Open Shop Scheduling Problem

Researchers propose a Transformer-based scheduling policy to solve the Open Shop Scheduling Problem, showing strong generalization from small to large instances.

AI/ML arXiv cs.AI

UP-NRPA: User Portrait based Nested Rollout Policy Adaptation for Planning with Large Language Models in Goal-oriented Dialogue Systems

UP-NRPA is introduced as an online framework for goal-oriented dialogue systems that adapts to user portraits without offline RL.

Other arXiv cs.AI

History of the Muddy Children Puzzle

A historical survey of the Muddy Children Puzzle and its influence on epistemic logic, including a new self-referential hats puzzle.

AI/ML arXiv cs.AI

Orchestra-o1: Omnimodal Agent Orchestration

Orchestra-o1 is presented as an omnimodal agent orchestration framework using DA-GRPO for efficient training of multi-agent collaboration.

AI/ML arXiv cs.AI

Hybrid Open-Ended Tri-Evolution Makes Better Deep Researcher

The HOTE framework utilizes hybrid-mode reinforcement learning to evolve proposer, solver, and judge modules for deep research agents.

AI/ML arXiv cs.AI

WorkBench Revisited: Workplace Agents Two Years On

A longitudinal study of the WorkBench benchmark shows significant improvements in agent capability and safety between 2024 and 2026.

AI/ML arXiv cs.AI

Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP

A comparison of Diff-in-Means and Iterative Nullspace Projection for steering refusal in open-weight chat models.