All Articles
16580 articles total
Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation
TRACER is a framework for clinical risk prediction that integrates severity-grounded knowledge graphs and RAG to improve outcomes in EHR analysis.
Using LLMs for Explainable, Data-Driven Insight Generation from Time Series
A domain-agnostic framework is proposed to generate explainable, data-driven natural language insights from time series forecasts, reducing LLM hallucinations.
Deep Reinforcement Learning to Master the Asymmetric Strategy of Baghchal
This study evaluates various Deep RL algorithms (DQN, PPO, MuZero) on the asymmetric board game Baghchal, finding MuZero to be the most effective.
Operational Hallucination and Safety Drift in AI Agents
The paper identifies 'Safety Drift' and 'Operational Hallucination' in AI agents and proposes an Action-Aware Supervision Layer to enforce architectural reliability.
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
AlayaWorld is an open-source interactive video world model capable of generating high-fps video environments from text, images, or video.
Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
A neuro-symbolic meta-policy is introduced for temporal knowledge-graph memory to improve decision-making under partial observability in RDF-based environments.
Original Apollo 11 Guidance Computer source code for command and lunar modules
The original source code for the Apollo 11 Guidance Computer (AGC), used for the first moon landing, is available for review.
Ask HN: US Equivalent of Anabin?
A community discussion on Hacker News seeking a US equivalent to the German academic recognition system, Anabin.
Probabilistic Concept-Aware Steering for Trustworthy LLM Inference
Researchers introduce the Probabilistic Concept-Aware Steering (PCS) framework to improve the interpretability and control of LLM inference via concept-driven steering vectors.
FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis
FindStatBench is a new execution benchmark designed to evaluate LLMs on combinatorial code synthesis, specifically mapping objects to integers or other objects.
When JSON Is Not Enough: Semantic Reliability of Schema-Constrained LLM Ordering Agents
OrderBench demonstrates that schema-constrained LLM outputs (like JSON) can still contain significant semantic errors, urging the need for domain-specific verification.
ProbSPARQL: Querying Knowledge Graphs with Multi-dimensional, Uncertain Numeric Data
ProbSPARQL is introduced as a SPARQL extension for querying knowledge graphs containing multi-dimensional, uncertain numeric sensor data.
Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII)
A position paper arguing that deepfake research is misaligned with the reality of non-consensual intimate imagery (AIG-NCII), focusing too much on authenticity detection.
MUX: Continuous Reasoning via Multiplexed Tokens
The MUX method proposes using multiplexed tokens in a latent space to achieve higher-bandwidth, more efficient continuous reasoning in LLMs.
State Compression in Two-Agent LLM Relays: A Closed-World Study of Constraint Preservation
A study on state compression in LLM agent relays shows that structured JSON hand-offs provide higher feasibility accuracy than narrative summarization.
Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models
Research shows that structured synthetic reasoning data can significantly improve arithmetic performance in small language models like Qwen3-0.6B/1.7B on consumer hardware.
Airglow browser lets users modify YouTube, Gmail and Spotify in real time
Airglow is a browser that enables real-time modification of popular web applications like YouTube, Gmail, and Spotify.
Kimi K3: second only to Fable 5 on AA-Briefcase
Kimi K3 model performance is reported as being second only to Fable 5 on the AA-Briefcase benchmark.
Ten ways a check passes while the thing it checks is broken
An exploration of ten common failure modes where a check passes despite the system being broken.
Integro-differential equations in angular stabilization of drone motion by distributed feedback control
A research paper proposing angular stabilization of drone motion using distributed feedback control and integro-differential equations.