Other arXiv cs.AI

Global Index on Responsible AI: 2026 Report

The 2026 Global Index on Responsible AI (GIRAI) report assesses how countries implement responsible AI governance and identifies gaps between policy and enforcement.

AI/ML arXiv cs.AI

Transcoders for Investigating Deception in Language Models

Uses transcoders to analyze deceptive behavior in language models, identifying specific features that contribute to dishonest outputs.

AI/ML arXiv cs.AI

CrimeNER Demo: Named-Entity Recognition in the Crime Domain

Presents CrimeNER Demo, a platform for extracting and classifying crime-related entities from documents using pretrained NER models.

AI/ML arXiv cs.AI

Reachability-Aware Pretraining for Efficient Target-Oriented Path Exploration in Temporal Knowledge Graph Reasoning

Introduces RAPTOR, a self-supervised pretraining method to improve reinforcement learning-based multi-hop reasoning in Temporal Knowledge Graphs.

Software Engineering arXiv cs.AI

Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control

Proposes Proof-or-Stop, a lifecycle control method for autonomous coding agents that requires verifiable evidence for state transitions.

AI/ML arXiv cs.AI

Contextualized Early Detection of Online Firestorms: A Sequential LLM-Based Approach

A sequential LLM-based approach for the early detection of online firestorms on social media platforms like Reddit.

Other Hacker News

EEG shows brain can simultaneous encode two speech streams

Research indicates that the human brain is capable of simultaneously encoding two separate streams of speech.

AI/ML arXiv cs.AI

Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology

Introduction of a new method for creating 3D image-text datasets for brain oncology and a corresponding VLM that outperforms 2D and 3D methods in MRI report generation.

AI/ML arXiv cs.AI

MathCoPilot: An Interactive System for Human-AI Symbiotic Paradigm of Mathematical Research

MathCoPilot is a human-in-the-loop system designed for mathematical research, allowing mathematicians to guide AI agents in formalizing and proving theorems in Lean 4.

AI/ML arXiv cs.AI

SportD: Can VLMs Physically Strategize?

The SportD benchmark evaluates whether Vision-Language Models (VLMs) can make strategically effective decisions in soccer, finding they often prefer lower-reward actions compared to pros.

AI/ML arXiv cs.AI

Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models

Action QFormer introduces a query-based interface to reorganize multimodal information into action-facing representations, significantly improving zero-shot sim-to-real navigation.

AI/ML arXiv cs.AI

Analytic Abduction: Causal Decomposition and Governed Commitment for Human--AI Coordination

Analytic Abduction is a framework for human-AI coordination that uses causal decomposition and governed commitment to avoid premature convergence in complex reasoning tasks.

AI/ML arXiv cs.AI

MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers

MCPEvol-Bench is a new benchmark that tests LLM agents' ability to adapt to evolving tool interfaces within Model Context Protocol (MCP) servers.

AI/ML arXiv cs.AI

TopoAgent: A Self-Evolving Topological Agent for Multimodal Scientific Reasoning

TopoAgent uses a self-evolving topological framework (DAG) to replace linear planning in multimodal scientific reasoning, reducing hallucinations and noise.

AI/ML arXiv cs.AI

SmartRAG: Native Graph-Based RAG for Mobile Device

SmartRAG is an on-device framework for mobile devices that uses a native graph-based RAG to achieve high-performance reasoning with a small quantized backbone.

AI/ML arXiv cs.AI

Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

Project Kaleidoscope provides a workflow for contextual, human-aligned evaluation of AI applications, integrating persona-based tests and reliability-gated scoring.

Other Hacker News

Starlink from 1984

A Hacker News discussion about a historical perspective or analogy relating Starlink to concepts from 1984.

AI/ML arXiv cs.AI

VLT: A Vision-Language-Time Series Multimodal Foundation Model for Industrial Intelligence

Introduction of VLT, a multimodal foundation model designed for industrial intelligence by bridging time-series data, frequency-spectrum visuals, and text.

AI/ML arXiv cs.AI

RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning

RetroAgent is introduced as an LLM agent that uses structured memory and chemistry tools to plan multi-step retrosynthesis for molecules.

AI/ML arXiv cs.AI

WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays

WrAFT is a modular automated writing evaluation system that provides scoring and feedback for argumentative essays using various LLMs.