All Articles
17547 articles total
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving
A comprehensive survey of KV cache management for LLM serving, classifying over 30 systems into five architectural archetypes.
Criterion-Conditional In-Context Learning: Evaluating Criterion-Shift Adaptation in Vision-Language Models
Introduction of Criterion-Conditional In-Context Learning (CC-ICL) and the CC-Bench benchmark to evaluate vision-language models' adaptability to shifting decision criteria.
An interactive explorer for Benford's Law across real datasets
An interactive explorer for visualizing Benford's Law across various real-world datasets.
Show HN: Fortress – a stealth Chromium so your agents stop getting blocked
Fortress is a stealth Chromium-based browser designed to prevent automation agents from being blocked by bot detection.
Honey, We Bought an AI Story
An opinion piece discussing the narrative and hype surrounding AI stories and marketing.
Structure and Interpretation of Computer Programs Video Lectures
Video lectures for the seminal 'Structure and Interpretation of Computer Programs' (SICP) course.
A 13th-Century Enumeration Algorithm, Ignored for 700 Years
Exploration of a 13th-century enumeration algorithm that remained largely ignored for 700 years.
The Hidden Water Geography of U.S. Hyperscale Data Centers in the AI Era
A study mapping the water consumption of U.S. hyperscale data centers, distinguishing between local cooling and regional electricity generation.
Not Every Sync Is Safe: Calibrated DiLoCo Scheduling for Shared AI Infrastructure
Introduction of WA-DiLoCo, a workload-aware controller for optimizing synchronization in fragmented AI training fleets to reduce SLO violations.
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences
DELTAVID is a framework and benchmark for improving fine-grained spatiotemporal perception in video-MLLMs using cross-video differences.
Double-Helix Active Geometry: LiDAR-Anchored Multi-View Depth with Selective Abstention
DH-Active is a lightweight, training-free geometry back-end for high-speed depth recovery from sparse LiDAR data on consumer devices.
Attention Dynamics in Diffusion Models: A Visual Analytics Framework for Human-AI Collaboration
A visual analytics framework for interpreting the evolution of attention dynamics and semantic structure in diffusion-based text-to-image models.
First Principles of Model Routing
A discussion on the first principles of model routing, focusing on how to efficiently direct queries to the most appropriate AI model.
We charge $10k a week to delete AI-generated code
A provocative claim or business model where a service charges $10k a week to remove AI-generated code from a codebase.
SiamixFormer: a fully-transformer Siamese network with temporal Fusion for accurate building detection and change detection in bi-temporal remote sensing images
Introduction of SiamixFormer, a transformer-based Siamese network for building and change detection in remote sensing images.
PotatoGANs: Utilizing Generative Adversarial Networks, Instance Segmentation, and Explainable AI for Enhanced Potato Disease Identification and Classification
PotatoGANs uses GANs and Explainable AI to improve the identification and classification of potato diseases through synthetic data augmentation.
Specific Domain Ontology Construction Using Large Language Models
Exploration of using LLMs like GPT-3.5 and GPT-4 to automatically construct specific domain ontologies, specifically for the Brazilian maritime territory.
Neural-Network Inverse Design of SRF Cavities and Transmons for Bosonic Quantum Computation
Use of deep neural networks for the inverse design of SRF cavities and transmons to optimize bosonic quantum computation hardware.
GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads
A technical report on optimizing GLM-5 serving parameters for the OpenClaw MaaS architecture to improve throughput and reduce latency for long-context agent workloads.
AutoResearch: An Execution-Grounded Multi-Agent Framework for Reliable Research Workflow Automation
AutoResearch is a multi-agent framework that automates research workflows with execution-grounding, including code repair and citation verification.