AI/ML arXiv cs.AI

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving

A comprehensive survey of KV cache management for LLM serving, classifying over 30 systems into five architectural archetypes.

AI/ML arXiv cs.AI

Criterion-Conditional In-Context Learning: Evaluating Criterion-Shift Adaptation in Vision-Language Models

Introduction of Criterion-Conditional In-Context Learning (CC-ICL) and the CC-Bench benchmark to evaluate vision-language models' adaptability to shifting decision criteria.

Other Hacker News

An interactive explorer for Benford's Law across real datasets

An interactive explorer for visualizing Benford's Law across various real-world datasets.

Software Engineering Hacker News

Show HN: Fortress – a stealth Chromium so your agents stop getting blocked

Fortress is a stealth Chromium-based browser designed to prevent automation agents from being blocked by bot detection.

AI/ML Hacker News

Honey, We Bought an AI Story

An opinion piece discussing the narrative and hype surrounding AI stories and marketing.

Software Engineering Hacker News

Structure and Interpretation of Computer Programs Video Lectures

Video lectures for the seminal 'Structure and Interpretation of Computer Programs' (SICP) course.

Software Engineering Hacker News

A 13th-Century Enumeration Algorithm, Ignored for 700 Years

Exploration of a 13th-century enumeration algorithm that remained largely ignored for 700 years.

Other arXiv cs.AI

The Hidden Water Geography of U.S. Hyperscale Data Centers in the AI Era

A study mapping the water consumption of U.S. hyperscale data centers, distinguishing between local cooling and regional electricity generation.

AI/ML arXiv cs.AI

Not Every Sync Is Safe: Calibrated DiLoCo Scheduling for Shared AI Infrastructure

Introduction of WA-DiLoCo, a workload-aware controller for optimizing synchronization in fragmented AI training fleets to reduce SLO violations.

AI/ML arXiv cs.AI

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences

DELTAVID is a framework and benchmark for improving fine-grained spatiotemporal perception in video-MLLMs using cross-video differences.

Hardware/Chips arXiv cs.AI

Double-Helix Active Geometry: LiDAR-Anchored Multi-View Depth with Selective Abstention

DH-Active is a lightweight, training-free geometry back-end for high-speed depth recovery from sparse LiDAR data on consumer devices.

AI/ML arXiv cs.AI

Attention Dynamics in Diffusion Models: A Visual Analytics Framework for Human-AI Collaboration

A visual analytics framework for interpreting the evolution of attention dynamics and semantic structure in diffusion-based text-to-image models.

AI/ML Hacker News

First Principles of Model Routing

A discussion on the first principles of model routing, focusing on how to efficiently direct queries to the most appropriate AI model.

Tech Business/VC Hacker News

We charge $10k a week to delete AI-generated code

A provocative claim or business model where a service charges $10k a week to remove AI-generated code from a codebase.

AI/ML arXiv cs.AI

SiamixFormer: a fully-transformer Siamese network with temporal Fusion for accurate building detection and change detection in bi-temporal remote sensing images

Introduction of SiamixFormer, a transformer-based Siamese network for building and change detection in remote sensing images.

AI/ML arXiv cs.AI

PotatoGANs: Utilizing Generative Adversarial Networks, Instance Segmentation, and Explainable AI for Enhanced Potato Disease Identification and Classification

PotatoGANs uses GANs and Explainable AI to improve the identification and classification of potato diseases through synthetic data augmentation.

AI/ML arXiv cs.AI

Specific Domain Ontology Construction Using Large Language Models

Exploration of using LLMs like GPT-3.5 and GPT-4 to automatically construct specific domain ontologies, specifically for the Brazilian maritime territory.

Hardware/Chips arXiv cs.AI

Neural-Network Inverse Design of SRF Cavities and Transmons for Bosonic Quantum Computation

Use of deep neural networks for the inverse design of SRF cavities and transmons to optimize bosonic quantum computation hardware.

AI/ML arXiv cs.AI

GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads

A technical report on optimizing GLM-5 serving parameters for the OpenClaw MaaS architecture to improve throughput and reduce latency for long-context agent workloads.

AI/ML arXiv cs.AI

AutoResearch: An Execution-Grounded Multi-Agent Framework for Reliable Research Workflow Automation

AutoResearch is a multi-agent framework that automates research workflows with execution-grounding, including code repair and citation verification.