All Articles
16086 articles total
ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection
ScratchSim provides a procedural synthetic data pipeline using BlenderProc to generate annotated data for surface scratch detection in industrial quality control.
SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence
SciFigAlign is a multimodal scorer that evaluates the quality of scientific figures by aligning them with manuscript evidence.
We shall dwell amidst wonder and glory for ever: On weird fiction
A discussion on the nature and appeal of weird fiction.
A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities
Research evaluates how coding agents comply with open-source contribution rules, finding that agents rarely retrieve rules proactively and struggle with bans.
AI as Friction for Reflection Support in Ideation
Proposes reframing AI in design ideation as a 'friction agent' to support reflection rather than just a tool for speeding up output.
Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
Introduces CostAda, a cost-calibrated adaptive controller for LLM discovery that optimizes search quality under fixed token budgets.
From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs
Explores trait-like internal representations in LLMs using SAE decomposition to link internal states with situational behaviors.
ReCo: Reweighting GRPO Against Distributional Concentration
Presents ReCo, a reweighting method for GRPO that prevents the model from concentrating on high-probability responses to preserve reasoning capacity.
Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents
Introduces TSDS, a framework for edge LLM agents that optimizes reasoning compute and defers to cloud models based on uncertainty.
Hearsay: Vision-Language Medical Diagnoses Without an Image
Reveals that VLMs often confabulate medical diagnoses based on patient demographics when the image is missing, showing systemic bias.
Human diversity fuels collective creativity that large language models cannot simulate or sustain
Demonstrates that AI ideation can homogenize creative output and fails to simulate the diversity produced by human groups.
Actions Have Consequences: Detecting Outcome Performativity using Intervention Testing
Formalizes Outcome Performativity A/B Detection (OPAB) to detect when predictions causally influence the outcomes they predict.
The End of an Era
A Hacker News discussion thread regarding 'The End of an Era', lacking specific content in the snippet.
Ruby Central's Destructive Legacy
A Hacker News discussion thread regarding Ruby Central's legacy and perceived destructive impacts.
Premier league bans gambling sponsors
News regarding the Premier League banning gambling sponsors.
D&D is getting World of Warcraft and Star Wars crossovers
Wizards of the Coast announces D&D crossovers with World of Warcraft and Star Wars.
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
Introduction of See2Think, a framework to evaluate whether multimodal LLMs actually utilize intermediate visual states during reasoning.
Journey Operators for Structured Multi-Axis Composition
A framework for multi-axis structure modeling and the introduction of JoFormer, relating journey operators to RoPE and state-space models.
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
SkillRise is presented as an agentic RL framework that allows LLM agents to evolve and reuse skills across related tasks via a curated skill document.
SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
SecRespond is a new benchmark for evaluating LLM agents' ability to perform post-compromise incident response in cybersecurity.