All Articles
17287 articles total
RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments
RetailBench is introduced as a simulation benchmark to evaluate the long-horizon autonomous decision-making and strategy stability of LLM agents in retail environments.
Preemption is GC for memory reordering (2019)
A discussion on how preemption in operating systems can be viewed as a form of garbage collection for memory reordering.
Meta removes controversial AI feature on Instagram after backlash
Meta has removed a controversial AI feature on Instagram following significant user backlash.
No, Flock isn’t threatening people for debating surveillance
Flock Safety is facing criticism for allegedly attempting to silence discussions about its surveillance technology via cease and desist letters.
Meta turns off the Instagram feature that let users make AI deepfakes of public accounts
Meta disables an Instagram AI feature that allowed users to create deepfakes of public accounts without permission.
A Practical Investigation of Training-free Relaxed Speculative Decoding
An investigation into training-free relaxed speculative decoding for accelerating LLM sampling, highlighting capability-speed trade-offs.
ProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation
Introduction of ProjAgent, a system for repository-level code generation using procedural similarity retrieval to better handle complex dependencies.
Pose-to-Biomechanics: Bridging 3D Human Pose Estimation and Biomechanical Attribute Prediction
BioModule is proposed as a lightweight plug-in for 3D pose estimators to predict biomechanical attributes for motion analysis.
Validity of LLMs as data annotators: AMALIA on authority
A study on the AMALIA-9B model's validity as a data annotator for European Portuguese, questioning its reliance on surface correlates.
Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph
Research demonstrating how applying graph algorithms to UMAP's internal kNN graph can enhance high-dimensional data sensemaking.
SLORR: Simple and Efficient In-Training Low-Rank Regularization
Introduction of SLORR, a stateless in-training low-rank regularization framework to improve neural network compressibility with minimal overhead.
Quantum error correction can constantly recalibrate a processor
Researchers are using reinforcement learning to constantly recalibrate quantum processors through error correction, improving stability.
Increased drone surveillance of illegal July 4th fireworks led to $100K fine
Increased drone surveillance by police and firefighters for illegal fireworks on July 4th led to significant fines.
When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability
This research audits the reliability of LLM-as-judge systems, finding that judge upgrades are not interchangeable and that stronger judges don't fully remove bias.
DocMaster: A Hierarchical Structure-Aware System for Document Analysis
DocMaster is a hierarchical structure-aware system for document analysis that preserves layout information for better filtering and QA.
VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval
VocaDet is a sample-driven open-vocabulary object detection framework that uses visual tokenization and vector databases for scalable recognition.
SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling
SMetric proposes a balanced session-centric scheduling approach for LLM serving agents to increase throughput and KV cache reuse.
When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities
The S2AE method improves mechanistic interpretability in vision-language models by enforcing concept consistency across modalities.
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing
UltraX is a function-calling refinement framework designed to improve the quality and efficiency of large-scale LLM pre-training data.
Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning
A new hierarchical machine teaching algorithm for reward learning is proposed to create autonomous agents robust across multiple environments.