All Articles
17670 articles total
What Drives Interactive Improvement from Feedback?
A study on how natural-language feedback improves AI agents, concluding that the student's ability to use feedback is a bigger bottleneck than the feedback itself.
Contrastive Reflection for Iterative Prompt Optimization
Introduction of Contrastive Reflection, an iterative prompt-optimization framework designed to improve agentic IR workflows through targeted prompt edits.
How Can AI Find My Model? A Model-Finding Experimental Study Considering Data Formats, Embeddings, and Retrieval Strategies
An experimental study on using AI and transformer-based embeddings to discover reusable simulation models via natural language queries.
BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation
Introduction of BayesBench, a suite of simulation environments used to evaluate how LLMs update beliefs in multi-turn evidence accumulation.
When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models
Research on 'LearnStop', a checkpoint stopper for reasoning models that predicts when to stop computation based on online features to save costs.
Beyond expert users: agents should help users construct preferences, not just elicit them
A proposal for CoPref and a benchmark called CoShop, arguing that agents should help users construct their preferences rather than just eliciting them.
Investigating Multi-Agent Deliberation in Law
Exploration of multi-agent deliberation (MAD) frameworks inspired by courtroom procedures to improve legal reasoning tasks in AI.
Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering
Introduction of HASTE, a hierarchical multi-agent system that allows ML engineering agents to transfer and reuse skills across different competitions.
RoPoLL: Robust Panel of LLM Judges
Presentation of RoPoLL, a robust panel of LLM judges that uses a geometric median estimator to reduce bias caused by corrupted judges.
Supersonic flight returning to US after half-century ban
The US is seeing a return of supersonic flight after a fifty-year ban.
Americans see their country's past, present and future
An article reflecting on the past, present, and future of the United States.
Manufactured Confidence: How Memory Consolidation Turns Hearsay into Confident Facts
Research demonstrates that LLM memory consolidation often converts tentative remarks into confident facts, creating a security vulnerability.
Deterministic Decisions for High-Stakes AI. A Zero-Egress Pipeline with the Deployability of RAG and the Accuracy of Machine Learning
A study reveals 'intervention bias' in zero-shot LLMs for educational advisory, proposing a supervised policy learning approach using Decision Transformers and XGBoost for better calibration.
Covering the Unseen: Information Demand Coverage Optimization for Retrieval-Augmented Generation
GeoRAG is introduced to optimize information demand coverage in RAG by treating context selection as a distribution optimization problem rather than simple ranking.
From Julia to Rust: a differentiable tensor stack for scientific computing
A discussion on transitioning a differentiable tensor stack for scientific computing from Julia to Rust for improved performance and safety.
Forestiere Underground Gardens
An article about the Forestiere Underground Gardens, which is not technical in nature.
Deriving the SVD (Single Value Decomposition) from scratch
A technical derivation of Single Value Decomposition (SVD) from scratch, providing a deep dive into the linear algebra behind it.
Scaling Laws, Carefully
A critical examination of scaling laws in AI, focusing on a more careful and nuanced approach to their application.
Structural Correctness
A discussion on the concept of structural correctness in software systems, emphasizing formal methods and reliability.
The “Father of the Internet” is finally retiring
Vinton Cerf, a co-creator of the internet's foundational protocols, is retiring from his role at Google.