AI/ML arXiv cs.AI

Copy-on-Write Scoring: Application-Specific Agent Evaluations

CoW Scoring uses a PostgreSQL-level Copy-on-Write mechanism to isolate agent writes, allowing for granular and inexpensive evaluation of LLM-based agents in real applications.

AI/ML arXiv cs.AI

Beyond scalar losses: calibrating segmentation models via gradient vector field surgery

The paper introduces a gradient vector field surgery technique to mitigate overconfidence and miscalibration in region-based loss functions for medical image segmentation.

Cybersecurity arXiv cs.AI

The Prover Is the Judge: Verified Security Software from AI Coding Agents in Ada/SPARK

A study demonstrates the use of AI coding agents to write and verify bare-metal security software in Ada/SPARK, using GNATprove to ensure functional correctness.

AI/ML arXiv cs.AI

Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution

GCA-Bench is introduced as a new benchmark for evaluating robotic grasping that requires multi-step reasoning and semantic understanding beyond simple visual detection.

AI/ML arXiv cs.AI

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

Research reveals 'value leakage,' where LLMs' answers are silently influenced by the values of their developers or the models themselves without disclosure.

AI/ML arXiv cs.AI

HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization

The HABIB_TAZ system uses mDeBERTa-v3 and synthetic training to disentangle formal logic from content plausibility in multilingual NLP tasks.

Software Engineering Hacker News

SQLite Is All You Need

A discussion on the utility and sufficiency of using SQLite for various application architectures, emphasizing its versatility.

Other Hacker News

Just got an AWS billing alert projecting my monthly cost at $140B

A user reports a massive, likely erroneous, AWS billing projection of $140 billion.

Other Hacker News

Ask HN: Any AWS billing issues known? Amazon forecast of 3 billion dollars

Hacker News community discussion regarding widespread or specific AWS billing anomalies and extreme forecasts.

Other The Verge

The war on ‘woke science’ comes for space research

Report on a proposal by the OMB to give political appointees more control over federal science grant funding, potentially impacting space research.

AI/ML arXiv cs.AI

Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection

Introduction of LIFT, a framework for accelerating VLA post-training by injecting reactive force data to improve robot manipulation in contact-rich environments.

AI/ML arXiv cs.AI

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning

MEMORA is presented as an Embodied Action Memory framework that allows robots to use egocentric videos for long-horizon reasoning and planning.

AI/ML arXiv cs.AI

Local Additive Feature Attribution: A Mathematical Taxonomy and Reporting Checklist

A mathematical taxonomy and reporting checklist for local additive feature attribution methods used in explainable AI (XAI).

AI/ML arXiv cs.AI

ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs

Introduction of ToolAlignBench, a benchmark for investigating alignment conflicts in tool-calling LLMs, particularly when safety training conflicts with deployment instructions.

AI/ML arXiv cs.AI

Assessing AI in Introductory Physics Problem Solving

Evaluation of OpenAI's o4-mini model's ability to solve introductory physics problems, noting performance drops in multimodal and high-difficulty tasks.

AI/ML arXiv cs.AI

Towards a Unified Multidimensional Explainability Metric: Evaluating Trustworthiness in AI Models

Proposed multidimensional framework for evaluating XAI methods like LIME and SHAP based on fidelity, simplicity, and stability.

Other Hacker News

I Owe My Life to the Commodore 64

A personal reflection on the lifelong impact of the Commodore 64 computer.

Tech Business/VC Hacker News

AWS: Inaccurate Estimated Billing Data - $1.7 BILLION

Discussion regarding significant inaccuracies in AWS estimated billing data, totaling $1.7 billion.

AI/ML arXiv cs.AI

The Cost and Network Limits of Space-Based AI Compute

Research analyzing the cost-effectiveness and network limitations of deploying AI data centers in low-Earth orbit compared to terrestrial facilities.

AI/ML arXiv cs.AI

RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences

Introduction of RENEW, a framework that uses human preferences and epistemic uncertainty to repair model exploitation in offline reinforcement learning world models.