AI/ML arXiv cs.AI

RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments

RetailBench is introduced as a simulation benchmark to evaluate the long-horizon autonomous decision-making and strategy stability of LLM agents in retail environments.

Software Engineering Hacker News

Preemption is GC for memory reordering (2019)

A discussion on how preemption in operating systems can be viewed as a form of garbage collection for memory reordering.

Tech Business/VC TechCrunch

Meta removes controversial AI feature on Instagram after backlash

Meta has removed a controversial AI feature on Instagram following significant user backlash.

Other The Verge

No, Flock isn’t threatening people for debating surveillance

Flock Safety is facing criticism for allegedly attempting to silence discussions about its surveillance technology via cease and desist letters.

AI/ML The Verge

Meta turns off the Instagram feature that let users make AI deepfakes of public accounts

Meta disables an Instagram AI feature that allowed users to create deepfakes of public accounts without permission.

AI/ML arXiv cs.AI

A Practical Investigation of Training-free Relaxed Speculative Decoding

An investigation into training-free relaxed speculative decoding for accelerating LLM sampling, highlighting capability-speed trade-offs.

AI/ML arXiv cs.AI

ProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation

Introduction of ProjAgent, a system for repository-level code generation using procedural similarity retrieval to better handle complex dependencies.

AI/ML arXiv cs.AI

Pose-to-Biomechanics: Bridging 3D Human Pose Estimation and Biomechanical Attribute Prediction

BioModule is proposed as a lightweight plug-in for 3D pose estimators to predict biomechanical attributes for motion analysis.

AI/ML arXiv cs.AI

Validity of LLMs as data annotators: AMALIA on authority

A study on the AMALIA-9B model's validity as a data annotator for European Portuguese, questioning its reliance on surface correlates.

AI/ML arXiv cs.AI

Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph

Research demonstrating how applying graph algorithms to UMAP's internal kNN graph can enhance high-dimensional data sensemaking.

AI/ML arXiv cs.AI

SLORR: Simple and Efficient In-Training Low-Rank Regularization

Introduction of SLORR, a stateless in-training low-rank regularization framework to improve neural network compressibility with minimal overhead.

Hardware/Chips Ars Technica

Quantum error correction can constantly recalibrate a processor

Researchers are using reinforcement learning to constantly recalibrate quantum processors through error correction, improving stability.

Other Ars Technica

Increased drone surveillance of illegal July 4th fireworks led to $100K fine

Increased drone surveillance by police and firefighters for illegal fireworks on July 4th led to significant fines.

AI/ML arXiv cs.AI

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

This research audits the reliability of LLM-as-judge systems, finding that judge upgrades are not interchangeable and that stronger judges don't fully remove bias.

AI/ML arXiv cs.AI

DocMaster: A Hierarchical Structure-Aware System for Document Analysis

DocMaster is a hierarchical structure-aware system for document analysis that preserves layout information for better filtering and QA.

AI/ML arXiv cs.AI

VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval

VocaDet is a sample-driven open-vocabulary object detection framework that uses visual tokenization and vector databases for scalable recognition.

AI/ML arXiv cs.AI

SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling

SMetric proposes a balanced session-centric scheduling approach for LLM serving agents to increase throughput and KV cache reuse.

AI/ML arXiv cs.AI

When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities

The S2AE method improves mechanistic interpretability in vision-language models by enforcing concept consistency across modalities.

AI/ML arXiv cs.AI

UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing

UltraX is a function-calling refinement framework designed to improve the quality and efficiency of large-scale LLM pre-training data.

AI/ML arXiv cs.AI

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning

A new hierarchical machine teaching algorithm for reward learning is proposed to create autonomous agents robust across multiple environments.