All Articles
16900 articles total
Global Index on Responsible AI: 2026 Report
The 2026 Global Index on Responsible AI (GIRAI) report assesses how countries implement responsible AI governance and identifies gaps between policy and enforcement.
Transcoders for Investigating Deception in Language Models
Uses transcoders to analyze deceptive behavior in language models, identifying specific features that contribute to dishonest outputs.
CrimeNER Demo: Named-Entity Recognition in the Crime Domain
Presents CrimeNER Demo, a platform for extracting and classifying crime-related entities from documents using pretrained NER models.
Reachability-Aware Pretraining for Efficient Target-Oriented Path Exploration in Temporal Knowledge Graph Reasoning
Introduces RAPTOR, a self-supervised pretraining method to improve reinforcement learning-based multi-hop reasoning in Temporal Knowledge Graphs.
Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control
Proposes Proof-or-Stop, a lifecycle control method for autonomous coding agents that requires verifiable evidence for state transitions.
Contextualized Early Detection of Online Firestorms: A Sequential LLM-Based Approach
A sequential LLM-based approach for the early detection of online firestorms on social media platforms like Reddit.
EEG shows brain can simultaneous encode two speech streams
Research indicates that the human brain is capable of simultaneously encoding two separate streams of speech.
Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology
Introduction of a new method for creating 3D image-text datasets for brain oncology and a corresponding VLM that outperforms 2D and 3D methods in MRI report generation.
MathCoPilot: An Interactive System for Human-AI Symbiotic Paradigm of Mathematical Research
MathCoPilot is a human-in-the-loop system designed for mathematical research, allowing mathematicians to guide AI agents in formalizing and proving theorems in Lean 4.
SportD: Can VLMs Physically Strategize?
The SportD benchmark evaluates whether Vision-Language Models (VLMs) can make strategically effective decisions in soccer, finding they often prefer lower-reward actions compared to pros.
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models
Action QFormer introduces a query-based interface to reorganize multimodal information into action-facing representations, significantly improving zero-shot sim-to-real navigation.
Analytic Abduction: Causal Decomposition and Governed Commitment for Human--AI Coordination
Analytic Abduction is a framework for human-AI coordination that uses causal decomposition and governed commitment to avoid premature convergence in complex reasoning tasks.
MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
MCPEvol-Bench is a new benchmark that tests LLM agents' ability to adapt to evolving tool interfaces within Model Context Protocol (MCP) servers.
TopoAgent: A Self-Evolving Topological Agent for Multimodal Scientific Reasoning
TopoAgent uses a self-evolving topological framework (DAG) to replace linear planning in multimodal scientific reasoning, reducing hallucinations and noise.
SmartRAG: Native Graph-Based RAG for Mobile Device
SmartRAG is an on-device framework for mobile devices that uses a native graph-based RAG to achieve high-performance reasoning with a small quantized backbone.
Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications
Project Kaleidoscope provides a workflow for contextual, human-aligned evaluation of AI applications, integrating persona-based tests and reliability-gated scoring.
Starlink from 1984
A Hacker News discussion about a historical perspective or analogy relating Starlink to concepts from 1984.
VLT: A Vision-Language-Time Series Multimodal Foundation Model for Industrial Intelligence
Introduction of VLT, a multimodal foundation model designed for industrial intelligence by bridging time-series data, frequency-spectrum visuals, and text.
RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning
RetroAgent is introduced as an LLM agent that uses structured memory and chemistry tools to plan multi-step retrosynthesis for molecules.
WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays
WrAFT is a modular automated writing evaluation system that provides scoring and feedback for argumentative essays using various LLMs.