AI/ML arXiv cs.AI

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)

The study proposes a lightweight graph-based RAG system to reduce hallucinations and improve factual correctness in complex question answering tasks.

AI/ML arXiv cs.AI

Will the Agent Recuse, and Will It Stop? Measuring LLM-Agent Compliance with In-Band Governance Signals at the Access Door and Mid-Flight

Research on 'Recuse Signals' for LLM agents shows that while agents may follow access-time directives, they largely ignore in-band signals to stop mid-flight tasks.

AI/ML AI News

Nvidia bets physical AI can solve healthcare robotics’ data problem

Nvidia is promoting 'Physical AI' and its Medical Physics Simulation framework to help healthcare robots learn through embodied experience.

Tech Business/VC AI News

AMD to invest up to $5 billion in Anthropic under AI infrastructure deal

AMD is investing up to $5 billion in Anthropic as part of a massive infrastructure deal to deploy MI450-series accelerators.

Hardware/Chips ServeTheHome

AMD Advancing AI 2026 Keynote Live Coverage

Live coverage of AMD's Advancing AI 2026 keynote, focusing on upcoming hardware and software technologies.

Hardware/Chips ServeTheHome

Diving Deeper on NVIDIA’s Vera CPU: New Architectural Details and SPEC CPU 2026 Benchmarks

Nvidia releases detailed architectural specs and SPEC CPU 2026 benchmarks for its upcoming Vera server CPU and Olympus core.

Other Hacker News

Fractl.art – a fractal generator with 8 octillion different patterns

A fractal generator capable of producing 8 octillion unique patterns.

Cybersecurity TechCrunch

How AI guardrails are impeding the work of offensive cybersecurity researchers

Cybersecurity researchers report that AI safety guardrails from OpenAI and Anthropic are hindering their ability to find and exploit vulnerabilities.

AI/ML arXiv cs.AI

Measuring LLM Trust Allocation Across Conflicting Software Artifacts

Introduction of TRACE, a method to evaluate how LLMs prioritize conflicting information between code, documentation, and tests, revealing a bias toward trusting documentation over implementation.

AI/ML arXiv cs.AI

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution

A new 'vocabulary dropout' technique for co-evolutionary self-play in LLMs to maintain curriculum diversity and improve mathematical reasoning performance.

AI/ML arXiv cs.AI

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models

Research on self-preference bias (SPB) in LLM-as-a-judge evaluations, finding that judges favor their own outputs even with objective rubrics.

AI/ML arXiv cs.AI

A Unified Survival Benchmark for Temporal Dropout Risk Prediction in Learning Analytics

A study introducing a survival-oriented benchmark for predicting student dropout in learning analytics using the OULAD dataset.

AI/ML arXiv cs.AI

An Auditable Policy-Simulation Framework for Student Dropout in Intervention-Free Data

A proposed framework for simulating counterfactual policies to address student dropout in higher education using LMS engagement data.

AI/ML arXiv cs.AI

Internal Knowledge Without External Expression: Probing the Generalization Boundary of a Classical Chinese Language Model

Investigation into Classical Chinese language models, finding that while they can internally encode facts, they fail to express uncertainty (the 'humility paradox') without RLHF.

AI/ML arXiv cs.AI

Generative Augmented Inference of LLM-generated Data for Market Research: Theory and Empirical Evidence

The Generative Augmented Inference (GAI) framework improves market research estimation by using LLM-generated data as informative features rather than direct proxies.

AI/ML arXiv cs.AI

Information Aggregation with AI Agents

An experiment on AI agents' ability to aggregate private information through trading in prediction markets, showing limitations in complex reasoning about others.

AI/ML arXiv cs.AI

A Sheaf-Theoretic and Topological Perspective on Complex Network Modeling and Attention Mechanisms in Graph Neural Models

A research paper proposing a cellular sheaf theoretic framework to analyze feature diffusion and aggregation in graph neural networks.

AI/ML arXiv cs.AI

In-Run Data Shapley for Adam Optimizer

Introduces Adam-Aware In-Run Data Shapley, a method for efficient and accurate data attribution in machine learning models using the Adam optimizer.

AI/ML arXiv cs.AI

LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs

Presents LatentLens, a method for interpreting visual tokens in vision-language models by mapping latent representations to natural language descriptions.

Software Engineering arXiv cs.AI

AgentCgroup: Understanding and Controlling OS Resources of AI Agents

Presents AgentCgroup, an eBPF-based resource controller designed to manage OS-level resources for AI agents in multi-tenant environments.