All Articles
16483 articles total
Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)
The study proposes a lightweight graph-based RAG system to reduce hallucinations and improve factual correctness in complex question answering tasks.
Will the Agent Recuse, and Will It Stop? Measuring LLM-Agent Compliance with In-Band Governance Signals at the Access Door and Mid-Flight
Research on 'Recuse Signals' for LLM agents shows that while agents may follow access-time directives, they largely ignore in-band signals to stop mid-flight tasks.
Nvidia bets physical AI can solve healthcare robotics’ data problem
Nvidia is promoting 'Physical AI' and its Medical Physics Simulation framework to help healthcare robots learn through embodied experience.
AMD to invest up to $5 billion in Anthropic under AI infrastructure deal
AMD is investing up to $5 billion in Anthropic as part of a massive infrastructure deal to deploy MI450-series accelerators.
AMD Advancing AI 2026 Keynote Live Coverage
Live coverage of AMD's Advancing AI 2026 keynote, focusing on upcoming hardware and software technologies.
Diving Deeper on NVIDIA’s Vera CPU: New Architectural Details and SPEC CPU 2026 Benchmarks
Nvidia releases detailed architectural specs and SPEC CPU 2026 benchmarks for its upcoming Vera server CPU and Olympus core.
Fractl.art – a fractal generator with 8 octillion different patterns
A fractal generator capable of producing 8 octillion unique patterns.
How AI guardrails are impeding the work of offensive cybersecurity researchers
Cybersecurity researchers report that AI safety guardrails from OpenAI and Anthropic are hindering their ability to find and exploit vulnerabilities.
Measuring LLM Trust Allocation Across Conflicting Software Artifacts
Introduction of TRACE, a method to evaluate how LLMs prioritize conflicting information between code, documentation, and tests, revealing a bias toward trusting documentation over implementation.
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
A new 'vocabulary dropout' technique for co-evolutionary self-play in LLMs to maintain curriculum diversity and improve mathematical reasoning performance.
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
Research on self-preference bias (SPB) in LLM-as-a-judge evaluations, finding that judges favor their own outputs even with objective rubrics.
A Unified Survival Benchmark for Temporal Dropout Risk Prediction in Learning Analytics
A study introducing a survival-oriented benchmark for predicting student dropout in learning analytics using the OULAD dataset.
An Auditable Policy-Simulation Framework for Student Dropout in Intervention-Free Data
A proposed framework for simulating counterfactual policies to address student dropout in higher education using LMS engagement data.
Internal Knowledge Without External Expression: Probing the Generalization Boundary of a Classical Chinese Language Model
Investigation into Classical Chinese language models, finding that while they can internally encode facts, they fail to express uncertainty (the 'humility paradox') without RLHF.
Generative Augmented Inference of LLM-generated Data for Market Research: Theory and Empirical Evidence
The Generative Augmented Inference (GAI) framework improves market research estimation by using LLM-generated data as informative features rather than direct proxies.
Information Aggregation with AI Agents
An experiment on AI agents' ability to aggregate private information through trading in prediction markets, showing limitations in complex reasoning about others.
A Sheaf-Theoretic and Topological Perspective on Complex Network Modeling and Attention Mechanisms in Graph Neural Models
A research paper proposing a cellular sheaf theoretic framework to analyze feature diffusion and aggregation in graph neural networks.
In-Run Data Shapley for Adam Optimizer
Introduces Adam-Aware In-Run Data Shapley, a method for efficient and accurate data attribution in machine learning models using the Adam optimizer.
LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
Presents LatentLens, a method for interpreting visual tokens in vision-language models by mapping latent representations to natural language descriptions.
AgentCgroup: Understanding and Controlling OS Resources of AI Agents
Presents AgentCgroup, an eBPF-based resource controller designed to manage OS-level resources for AI agents in multi-tenant environments.