AI/ML arXiv cs.AI

ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection

ScratchSim provides a procedural synthetic data pipeline using BlenderProc to generate annotated data for surface scratch detection in industrial quality control.

AI/ML arXiv cs.AI

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence

SciFigAlign is a multimodal scorer that evaluates the quality of scientific figures by aligning them with manuscript evidence.

Other Hacker News

We shall dwell amidst wonder and glory for ever: On weird fiction

A discussion on the nature and appeal of weird fiction.

AI/ML arXiv cs.AI

A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities

Research evaluates how coding agents comply with open-source contribution rules, finding that agents rarely retrieve rules proactively and struggle with bans.

AI/ML arXiv cs.AI

AI as Friction for Reflection Support in Ideation

Proposes reframing AI in design ideation as a 'friction agent' to support reflection rather than just a tool for speeding up output.

AI/ML arXiv cs.AI

Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility

Introduces CostAda, a cost-calibrated adaptive controller for LLM discovery that optimizes search quality under fixed token budgets.

AI/ML arXiv cs.AI

From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

Explores trait-like internal representations in LLMs using SAE decomposition to link internal states with situational behaviors.

AI/ML arXiv cs.AI

ReCo: Reweighting GRPO Against Distributional Concentration

Presents ReCo, a reweighting method for GRPO that prevents the model from concentrating on high-probability responses to preserve reasoning capacity.

AI/ML arXiv cs.AI

Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents

Introduces TSDS, a framework for edge LLM agents that optimizes reasoning compute and defers to cloud models based on uncertainty.

AI/ML arXiv cs.AI

Hearsay: Vision-Language Medical Diagnoses Without an Image

Reveals that VLMs often confabulate medical diagnoses based on patient demographics when the image is missing, showing systemic bias.

AI/ML arXiv cs.AI

Human diversity fuels collective creativity that large language models cannot simulate or sustain

Demonstrates that AI ideation can homogenize creative output and fails to simulate the diversity produced by human groups.

AI/ML arXiv cs.AI

Actions Have Consequences: Detecting Outcome Performativity using Intervention Testing

Formalizes Outcome Performativity A/B Detection (OPAB) to detect when predictions causally influence the outcomes they predict.

Other Hacker News

The End of an Era

A Hacker News discussion thread regarding 'The End of an Era', lacking specific content in the snippet.

Open Source Hacker News

Ruby Central's Destructive Legacy

A Hacker News discussion thread regarding Ruby Central's legacy and perceived destructive impacts.

Other Hacker News

Premier league bans gambling sponsors

News regarding the Premier League banning gambling sponsors.

Other The Verge

D&D is getting World of Warcraft and Star Wars crossovers

Wizards of the Coast announces D&D crossovers with World of Warcraft and Star Wars.

AI/ML arXiv cs.AI

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Introduction of See2Think, a framework to evaluate whether multimodal LLMs actually utilize intermediate visual states during reasoning.

AI/ML arXiv cs.AI

Journey Operators for Structured Multi-Axis Composition

A framework for multi-axis structure modeling and the introduction of JoFormer, relating journey operators to RoPE and state-space models.

AI/ML arXiv cs.AI

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

SkillRise is presented as an agentic RL framework that allows LLM agents to evolve and reuse skills across related tasks via a curated skill document.

Cybersecurity arXiv cs.AI

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

SecRespond is a new benchmark for evaluating LLM agents' ability to perform post-compromise incident response in cybersecurity.