AI/ML arXiv cs.AI

The Large Cancer Assistant (LCA): A Model-Agnostic Orchestration Framework for Scalable Clinical Decision Support in Oncology

Introduction of the Large Cancer Assistant (LCA), a model-agnostic orchestration framework for scalable clinical decision support in oncology.

AI/ML arXiv cs.AI

Rethinking Indic AI from a Lens of Cultural Heritage Preservation

An exploration of Indic AI through the lens of cultural heritage preservation and the proposal of 'Culture Sensing' for more inclusive NLP models.

AI/ML arXiv cs.AI

Lingering Authority: Revocable Resource-and-Effect Capabilities for Coding Agents

Introduction of PORTICO, a reference monitor for revocable capabilities to prevent 'lingering authority' in coding agents.

AI/ML arXiv cs.AI

KVpop -- Key-Value Cache Compression with Predictive Online Pruning

Introduction of KVpop, a KV cache compression method using predictive online pruning to maintain performance with reduced memory costs.

Cybersecurity arXiv cs.AI

Proof of Execution: Runtime Verification for Governed AI Agent Actions

Proposed 'Proof of Execution' (PoE) framework to provide runtime verification and tamper-evident history for governed AI agent actions.

AI/ML arXiv cs.AI

Benchmarking KV-Cache Optimizations across Task Quality and System Performance for Long-Context Serving

A benchmarking study of various KV-cache optimization techniques (KIVI, TurboQuant, SnapKV, CaM) for long-context LLM serving.

AI/ML arXiv cs.AI

Catalyst Papers in Artificial Intelligence Research: A Landscape on ICLR from 2017 to 2025

A large-scale analysis of ICLR papers (2017-2025) identifying 'catalyst' papers and finding that peer review scores correlate poorly with future disruptiveness.

Other arXiv cs.AI

AI tools in Arab University English classrooms: Looking back and forward

A synthesis of empirical research on AI tool usage for English language learning in Arab University classrooms.

Hardware/Chips Hacker News

Home made GPU escalated quickly [video]

A video showcasing a DIY attempt at building a homemade GPU, highlighting the technical challenges and iterative process.

AI/ML TechCrunch

Hot French startup ZML releases free product to speed inference across lots of AI chips

French startup ZML has released LLMD, a software tool designed to optimize and speed up AI inference across various chip architectures.

Hardware/Chips The Verge

Samsung will launch its new wide foldable on July 22nd

Samsung announces a new Galaxy Unpacked event on July 22nd, expected to debut a new wide foldable phone format.

AI/ML arXiv cs.AI

Multi-Agent Deep Reinforcement Learning for Multi Objective Battery Management in Dairy Farms

Researchers propose a multi-agent Deep Reinforcement Learning system to optimize battery management and renewable energy integration in Irish dairy farms.

AI/ML arXiv cs.AI

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade

A new method called Recall-Controlled Probe Cascade allows LLM agents to abort failing trajectories early by analyzing internal hidden activations, saving significant compute.

AI/ML arXiv cs.AI

RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models

The RMISC corpus is introduced as a large-scale, real-world multivariate time series archive to improve the pretraining and generalization of time series foundation models.

AI/ML arXiv cs.AI

FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games

FootsiesGym is an open-source benchmark environment for two-player, zero-sum, imperfect-information games based on a 2D fighting game.

AI/ML arXiv cs.AI

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

FreqDepthKV is introduced as a cache compression method for long-context LLM inference that uses frequency-guided depth sharing to reduce memory and increase throughput.

AI/ML arXiv cs.AI

Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment

VAORA is a novel reward design for Vision-Language Models (VLMs) that aligns reasoning with visual outcomes to reduce hallucinations in physical reasoning tasks.

AI/ML arXiv cs.AI

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression

DepthWeave-KV proposes a token-adaptive cross-layer residual factorization method to achieve up to 8.3x KV cache memory reduction for long-context LLM inference.

Cybersecurity Ars Technica

Hackers can use 9 of the most popular AI tools to assemble massive botnets

Researchers have identified a new vulnerability called 'HalluSquatting' where hackers can leverage LLM hallucinations to trick AI tools into assembling massive botnets.

AI/ML arXiv cs.AI

Task Decomposition-Guided Reranking for Adaptive Agent Skill Retrieval

The SkillReranker framework improves AI agent skill selection by using semantic decomposition and a directed acyclic execution graph for adaptive reranking.