AI/ML arXiv cs.AI

Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent

The authors present Atrex-Bench for evaluating LLM-generated GPU kernels and introduce Atrex-Kernel-Agent (AKA) to optimize those kernels for production workloads.

AI/ML arXiv cs.AI

Towards an Intention Abstraction Layer for Autonomous Industrial Systems

The Intention Abstraction Layer (IAL) is proposed as middleware to parse human goals into structured intentions for autonomous industrial systems to avoid conflicts.

AI/ML arXiv cs.AI

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation

Seer is a training-free framework that accelerates Diffusion MLLMs by detecting semantic boundaries at the first step to truncate redundant computation.

Cybersecurity arXiv cs.AI

Democratizing Agent Deployment Safety: A Structural Monitoring Approach

The paper introduces an Information Flow Graph (IFG) monitor to detect security regressions and sabotage in AI software development agents.

AI/ML arXiv cs.AI

Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

Alipay-PIBench is a new benchmark for evaluating the ability of coding agents to perform realistic payment integration tasks.

AI/ML arXiv cs.AI

Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments

A study investigating collective spatial learning and network-efficiency effects among groups of LLM agents in simulated social experiments.

Other Hacker News

Pebble Mega Update – July 2026

A July 2026 update regarding the Pebble smartwatch ecosystem.

AI/ML Hacker News

UIUC AI Teaching Assistant

Discussion of an AI Teaching Assistant developed at UIUC.

AI/ML arXiv cs.AI

Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration

Research exploring how the design of evaluation instruments can significantly bias measured honesty in language models.

AI/ML arXiv cs.AI

Reward-Free Evolving Agents via Pairwise Validator

Introduces a pairwise validator method to replace scalar rewards in self-evolving agentic loops, reducing the cost of reward design.

AI/ML arXiv cs.AI

CausalGraphX: A Counterfactual Graph Neural Network Framework for Explainable Systemic Risk Assessment

Presents CausalGraphX, a framework combining GNNs and counterfactual reasoning for explainable systemic risk assessment in financial networks.

AI/ML arXiv cs.AI

Per-Token Fixed-Point Convergence in Depth-Recurrent Transformers

Analysis of depth-recurrent transformers showing per-token fixed-point convergence, enabling a training-free early-exit rule to reduce computation.

AI/ML arXiv cs.AI

Tactile: Giving Computer-Using Agents Hands and Feet

Introduces Tactile, an open-source tool layer that provides agents with a reliable, semantic interface for desktop automation beyond simple coordinates.

AI/ML arXiv cs.AI

Step-Level Preference Learning for Generative Agents in Social Simulations

Proposes step-level preference learning for generative agents in social simulations to improve fidelity and interaction quality.

AI/ML arXiv cs.AI

SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation

Introduces SAGA, a training-free framework for schema-aware grounding in agentic Text-to-SPARQL generation to reduce 'type-blind' errors.

AI/ML arXiv cs.AI

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions

Presents CEDI, a framework for the contextualized, multi-turn evaluation of Vision Language Models to better identify real-world hallucinations.

AI/ML Hacker News

What loss.backward() actually does

A discussion on the internal mechanics and mathematical operation of the loss.backward() function in deep learning frameworks.

AI/ML arXiv cs.AI

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation

Proposes an automated agentic red-teaming framework to synthesize difficult adversarial examples for improving the robustness of Multimodal LLMs.

AI/ML arXiv cs.AI

AI Agents Do Not Fail Alone:The Context Fails First

Presents ProofAgent-Harness, an open-source infrastructure for evaluating AI agent reliability by measuring context-engineering quality.

Other arXiv cs.AI

Measuring How Students Rely on Generative AI in Academic Writing: Development and Multi-Source Validation of the Generative AI Reliance Types Scale (GenAI-RTS)

Develops the GenAI-RTS scale to measure four types of student reliance on generative AI in academic writing.