AI/ML arXiv cs.AI

StrideDiffusion: Accelerating Diffusion Models for Time-series Generation

StrideDiffusion introduces a training-free spectral-aware sampler that significantly accelerates diffusion models for time-series generation.

AI/ML arXiv cs.AI

CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning

CMI-Mem is a lightweight memory manager for AI agents using RL with a reward hybrid of QA correctness and Conditional Mutual Information.

AI/ML arXiv cs.AI

AI-Driven Multi-Hop Relay Selection for Smart Urban NR-V2X Networks via Learning-to-Optimize Graph Neural Networks

A GNN-based Learning-to-Optimize framework for real-time multi-hop relay selection in urban NR-V2X networks to reduce latency and improve connectivity.

AI/ML arXiv cs.AI

KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback

KeySI is an interaction framework that allows users to tune text embeddings using keyword-based concept specification rather than document-level feedback.

AI/ML arXiv cs.AI

WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms

WaveformQA is an open-source benchmark for testing the temporal reasoning capabilities of LLMs on digital waveform data.

AI/ML arXiv cs.AI

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

NVIDIA presents NOOA, a model-agnostic Python framework where AI agents are defined as Python objects, treating methods as actions and docstrings as prompts.

Other Hacker News

Flux 3

No specific details provided for the Flux 3 article beyond the title and source.

Other Hacker News

Visualizing the Artemis II Mission

The article discusses a visual representation of the upcoming Artemis II mission.

Tech Business/VC Hacker News

Stripe in talks to buy OpenRouter for ~10B

Reports on potential talks between Stripe and OpenRouter regarding a multi-billion dollar acquisition.

AI/ML arXiv cs.AI

CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits

Presents CANN Bench, an open benchmark for evaluating AI-generated operator kernels on Huawei's Ascend NPU.

AI/ML arXiv cs.AI

Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving

Investigates how changes in problem representation affect the mathematical reasoning performance of LLMs.

AI/ML arXiv cs.AI

Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models

Studies attention degradation in transformer models and its causal relationship to contextual retrieval.

AI/ML arXiv cs.AI

Autonomous disproofs of the sum-product conjecture over $\mathbb R$ with GPT-5.5 Pro

An agent built on GPT-5.5 Pro autonomously generates proofs disproving the sum-product conjecture over the real numbers.

AI/ML arXiv cs.AI

ConfidenceBench: Evaluating Confidence Calibration in Large Language Models

Presents ConfidenceBench, a benchmark for evaluating how well LLMs verbalize their confidence estimates.

AI/ML arXiv cs.AI

Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis

Discusses the unreliability of current citation-checking agents and proposes a new gold-anchored evaluation protocol and guard.

AI/ML arXiv cs.AI

PromptPack: Scaling LLM Annotation Agents for Online Recommendation

Presents PromptPack, a method for scaling LLM annotation agents via in-context batching to reduce costs and increase throughput.

Other Hacker News

For most US drivers, EVs offer emissions benefits and cost savings

An article discussing the emission benefits and cost savings associated with electric vehicles for most US drivers.

Other Hacker News

Show HN: A factory simulator game built on Spreadsheets

A developer showcases a factory simulator game created entirely using spreadsheets.

AI/ML arXiv cs.AI

MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation

MKEvolve is a modular multi-agent framework designed to iteratively co-evolve PyTorch modules and LLM-generated kernels for hardware accelerators.

AI/ML arXiv cs.AI

Inducing Comparability of Factorised Probability Distributions

A research paper proposing an extension scheme to allow for principled comparison between factorised probability distributions defined over non-identical variable sets.