All Articles
19212 articles total
Recurrent Reasoning on Symbolic Puzzles with Sequence Models
Introduction of RecurrReason, a difficulty-controlled benchmark for testing recurrent logic puzzles on sequence models to analyze reasoning robustness.
Do LLMs Reliably Identify Correct Information Units in Aphasic Discourse?
A study evaluating the ability of LLMs like Llama-3.1 and Qwen2.5 to identify Correct Information Units (CIUs) in aphasic discourse transcripts.
Artificial Intelligence Index Report 2026
The 2026 AI Index Report provides a comprehensive overview of AI's progress in reasoning, safety, economy, and its impact on science and medicine.
AI-Driven Framework for Adaptive Water Network Management with Proof-of-Concept Implementation: Addressing Non-Revenue Water in Jordan
A framework for adaptive water network management in Jordan combining EPANET, digital twins, and local LLMs (via Ollama) to reduce non-revenue water loss.
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
Introduction of RoboPIN and Pinned Chain-of-Thought, a paradigm that anchors reasoning steps to visual evidence to improve embodied AI grounding.
John Carmack on Fabrice Bellard
A discussion on Hacker News featuring perspectives from John Carmack regarding the work of Fabrice Bellard.
Chili peppers of the world: cultivars, species, and heat
A community discussion about the various cultivars and heat levels of chili peppers worldwide.
QoS-Aware Token Scheduling and Private Data Valuation for Multi-Modal Agentic Networks
Research on fair token allocation and private data valuation in decentralized multi-modal agentic networks using differentially private prototypes.
Do we have the knowledge we need? Rethinking human-AI decision-making in corporations
A position paper exploring how organizations should maintain knowledge for both human and AI accessibility and how to allocate agency between them.
Large Language Models as Optimizers: A Survey of Direct vs. Tool-Augmented Approaches and Their Performance Frontiers
A survey of LLMs as optimizers, analyzing direct, tool-augmented, and tool-creating optimization paradigms.
Your Agent Has a Genome: Sequence-Level Behavioral Analysis and Runtime Governance of LLM-Powered Autonomous Agents
Introduces Base Sequence Analysis for LLM agent behavior and Governor, a runtime intervention system that increases success rates and reduces token costs.
Agentic Retrieval and Reinforcement Learned Equation Chains: A Controlled Generation Framework for Complex and Novel Physics Word Problems
Introduces ARVRE, a framework using reinforcement learning and agentic RAG to generate mathematically valid and complex physics word problems.
Integrating Reasoning and Generalization in Text-to-SQL via Self-Enhanced Fine-Tuning
Presents CoTE-SQL, a method for enhancing Text-to-SQL generation using self-enhanced reasoning traces and error-aware revision.
NeuroSymbolic AI for Legal AI-TRISM: Trustworthy, Reliable, Interpretable, Safe Models
Proposed TRISM framework combining NeuroSymbolic AI with LLMs to improve trust, reliability, and interpretability in legal AI applications.
Towards Next-Generation Healthcare: A Survey of Medical Embodied AI for Perception, Decision-Making, and Action
A comprehensive survey on Medical Embodied AI, focusing on the integration of perception, decision-making, and action in clinical environments.
I Could've Rickrolled the FIFA World Cup. All I Needed Was My ID
A first-person account of a potential security vulnerability at the FIFA World Cup involving ID badge access.
Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds
Research on reward hacking in LLM agents using a text-based evaluation suite to show that proxy-reward failures resist standard mitigations.
Hierarchical Modeling of ICD Codes in EHR Foundation Models
A study on improving EHR foundation models by explicitly incorporating the hierarchical structure of ICD-10-CM diagnosis codes.
Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines
Proposed method for anytime-valid attribution in LLM evaluation pipelines to distinguish between product drift and judge model drift.
Towards End-to-End Automation of AI Research
Introduction of 'The AI Scientist', an end-to-end automated system that can generate research ideas, execute experiments, and write scientific manuscripts.