All Articles
17543 articles total
The Large Cancer Assistant (LCA): A Model-Agnostic Orchestration Framework for Scalable Clinical Decision Support in Oncology
Introduction of the Large Cancer Assistant (LCA), a model-agnostic orchestration framework for scalable clinical decision support in oncology.
Rethinking Indic AI from a Lens of Cultural Heritage Preservation
An exploration of Indic AI through the lens of cultural heritage preservation and the proposal of 'Culture Sensing' for more inclusive NLP models.
Lingering Authority: Revocable Resource-and-Effect Capabilities for Coding Agents
Introduction of PORTICO, a reference monitor for revocable capabilities to prevent 'lingering authority' in coding agents.
KVpop -- Key-Value Cache Compression with Predictive Online Pruning
Introduction of KVpop, a KV cache compression method using predictive online pruning to maintain performance with reduced memory costs.
Proof of Execution: Runtime Verification for Governed AI Agent Actions
Proposed 'Proof of Execution' (PoE) framework to provide runtime verification and tamper-evident history for governed AI agent actions.
Benchmarking KV-Cache Optimizations across Task Quality and System Performance for Long-Context Serving
A benchmarking study of various KV-cache optimization techniques (KIVI, TurboQuant, SnapKV, CaM) for long-context LLM serving.
Catalyst Papers in Artificial Intelligence Research: A Landscape on ICLR from 2017 to 2025
A large-scale analysis of ICLR papers (2017-2025) identifying 'catalyst' papers and finding that peer review scores correlate poorly with future disruptiveness.
AI tools in Arab University English classrooms: Looking back and forward
A synthesis of empirical research on AI tool usage for English language learning in Arab University classrooms.
Home made GPU escalated quickly [video]
A video showcasing a DIY attempt at building a homemade GPU, highlighting the technical challenges and iterative process.
Hot French startup ZML releases free product to speed inference across lots of AI chips
French startup ZML has released LLMD, a software tool designed to optimize and speed up AI inference across various chip architectures.
Samsung will launch its new wide foldable on July 22nd
Samsung announces a new Galaxy Unpacked event on July 22nd, expected to debut a new wide foldable phone format.
Multi-Agent Deep Reinforcement Learning for Multi Objective Battery Management in Dairy Farms
Researchers propose a multi-agent Deep Reinforcement Learning system to optimize battery management and renewable energy integration in Irish dairy farms.
Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade
A new method called Recall-Controlled Probe Cascade allows LLM agents to abort failing trajectories early by analyzing internal hidden activations, saving significant compute.
RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models
The RMISC corpus is introduced as a large-scale, real-world multivariate time series archive to improve the pretraining and generalization of time series foundation models.
FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games
FootsiesGym is an open-source benchmark environment for two-player, zero-sum, imperfect-information games based on a 2D fighting game.
FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference
FreqDepthKV is introduced as a cache compression method for long-context LLM inference that uses frequency-guided depth sharing to reduce memory and increase throughput.
Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment
VAORA is a novel reward design for Vision-Language Models (VLMs) that aligns reasoning with visual outcomes to reduce hallucinations in physical reasoning tasks.
DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression
DepthWeave-KV proposes a token-adaptive cross-layer residual factorization method to achieve up to 8.3x KV cache memory reduction for long-context LLM inference.
Hackers can use 9 of the most popular AI tools to assemble massive botnets
Researchers have identified a new vulnerability called 'HalluSquatting' where hackers can leverage LLM hallucinations to trick AI tools into assembling massive botnets.
Task Decomposition-Guided Reranking for Adaptive Agent Skill Retrieval
The SkillReranker framework improves AI agent skill selection by using semantic decomposition and a directed acyclic execution graph for adaptive reranking.