All Articles
16880 articles total
Copy-on-Write Scoring: Application-Specific Agent Evaluations
CoW Scoring uses a PostgreSQL-level Copy-on-Write mechanism to isolate agent writes, allowing for granular and inexpensive evaluation of LLM-based agents in real applications.
Beyond scalar losses: calibrating segmentation models via gradient vector field surgery
The paper introduces a gradient vector field surgery technique to mitigate overconfidence and miscalibration in region-based loss functions for medical image segmentation.
The Prover Is the Judge: Verified Security Software from AI Coding Agents in Ada/SPARK
A study demonstrates the use of AI coding agents to write and verify bare-metal security software in Ada/SPARK, using GNATprove to ensure functional correctness.
Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution
GCA-Bench is introduced as a new benchmark for evaluating robotic grasping that requires multi-step reasoning and semantic understanding beyond simple visual detection.
Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
Research reveals 'value leakage,' where LLMs' answers are silently influenced by the values of their developers or the models themselves without disclosure.
HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization
The HABIB_TAZ system uses mDeBERTa-v3 and synthetic training to disentangle formal logic from content plausibility in multilingual NLP tasks.
SQLite Is All You Need
A discussion on the utility and sufficiency of using SQLite for various application architectures, emphasizing its versatility.
Just got an AWS billing alert projecting my monthly cost at $140B
A user reports a massive, likely erroneous, AWS billing projection of $140 billion.
Ask HN: Any AWS billing issues known? Amazon forecast of 3 billion dollars
Hacker News community discussion regarding widespread or specific AWS billing anomalies and extreme forecasts.
The war on ‘woke science’ comes for space research
Report on a proposal by the OMB to give political appointees more control over federal science grant funding, potentially impacting space research.
Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection
Introduction of LIFT, a framework for accelerating VLA post-training by injecting reactive force data to improve robot manipulation in contact-rich environments.
MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning
MEMORA is presented as an Embodied Action Memory framework that allows robots to use egocentric videos for long-horizon reasoning and planning.
Local Additive Feature Attribution: A Mathematical Taxonomy and Reporting Checklist
A mathematical taxonomy and reporting checklist for local additive feature attribution methods used in explainable AI (XAI).
ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
Introduction of ToolAlignBench, a benchmark for investigating alignment conflicts in tool-calling LLMs, particularly when safety training conflicts with deployment instructions.
Assessing AI in Introductory Physics Problem Solving
Evaluation of OpenAI's o4-mini model's ability to solve introductory physics problems, noting performance drops in multimodal and high-difficulty tasks.
Towards a Unified Multidimensional Explainability Metric: Evaluating Trustworthiness in AI Models
Proposed multidimensional framework for evaluating XAI methods like LIME and SHAP based on fidelity, simplicity, and stability.
I Owe My Life to the Commodore 64
A personal reflection on the lifelong impact of the Commodore 64 computer.
AWS: Inaccurate Estimated Billing Data - $1.7 BILLION
Discussion regarding significant inaccuracies in AWS estimated billing data, totaling $1.7 billion.
The Cost and Network Limits of Space-Based AI Compute
Research analyzing the cost-effectiveness and network limitations of deploying AI data centers in low-Earth orbit compared to terrestrial facilities.
RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
Introduction of RENEW, a framework that uses human preferences and epistemic uncertainty to repair model exploitation in offline reinforcement learning world models.