All Articles
16483 articles total
StrideDiffusion: Accelerating Diffusion Models for Time-series Generation
StrideDiffusion introduces a training-free spectral-aware sampler that significantly accelerates diffusion models for time-series generation.
CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning
CMI-Mem is a lightweight memory manager for AI agents using RL with a reward hybrid of QA correctness and Conditional Mutual Information.
AI-Driven Multi-Hop Relay Selection for Smart Urban NR-V2X Networks via Learning-to-Optimize Graph Neural Networks
A GNN-based Learning-to-Optimize framework for real-time multi-hop relay selection in urban NR-V2X networks to reduce latency and improve connectivity.
KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback
KeySI is an interaction framework that allows users to tune text embeddings using keyword-based concept specification rather than document-level feedback.
WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms
WaveformQA is an open-source benchmark for testing the temporal reasoning capabilities of LLMs on digital waveform data.
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
NVIDIA presents NOOA, a model-agnostic Python framework where AI agents are defined as Python objects, treating methods as actions and docstrings as prompts.
Flux 3
No specific details provided for the Flux 3 article beyond the title and source.
Visualizing the Artemis II Mission
The article discusses a visual representation of the upcoming Artemis II mission.
Stripe in talks to buy OpenRouter for ~10B
Reports on potential talks between Stripe and OpenRouter regarding a multi-billion dollar acquisition.
CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits
Presents CANN Bench, an open benchmark for evaluating AI-generated operator kernels on Huawei's Ascend NPU.
Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
Investigates how changes in problem representation affect the mathematical reasoning performance of LLMs.
Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
Studies attention degradation in transformer models and its causal relationship to contextual retrieval.
Autonomous disproofs of the sum-product conjecture over $\mathbb R$ with GPT-5.5 Pro
An agent built on GPT-5.5 Pro autonomously generates proofs disproving the sum-product conjecture over the real numbers.
ConfidenceBench: Evaluating Confidence Calibration in Large Language Models
Presents ConfidenceBench, a benchmark for evaluating how well LLMs verbalize their confidence estimates.
Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis
Discusses the unreliability of current citation-checking agents and proposes a new gold-anchored evaluation protocol and guard.
PromptPack: Scaling LLM Annotation Agents for Online Recommendation
Presents PromptPack, a method for scaling LLM annotation agents via in-context batching to reduce costs and increase throughput.
For most US drivers, EVs offer emissions benefits and cost savings
An article discussing the emission benefits and cost savings associated with electric vehicles for most US drivers.
Show HN: A factory simulator game built on Spreadsheets
A developer showcases a factory simulator game created entirely using spreadsheets.
MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation
MKEvolve is a modular multi-agent framework designed to iteratively co-evolve PyTorch modules and LLM-generated kernels for hardware accelerators.
Inducing Comparability of Factorised Probability Distributions
A research paper proposing an extension scheme to allow for principled comparison between factorised probability distributions defined over non-identical variable sets.