AI/ML arXiv cs.AI

A Three-Layer Framework for AI in Scientific Discovery

A proposed three-layer framework for AI in scientific discovery, emphasizing 'model formation' via qualitative reasoning over simple search or execution.

AI/ML arXiv cs.AI

Replication in Visual Diffusion Models: A Survey and Outlook

A comprehensive survey on the phenomenon of replication (memorization) in visual diffusion models, covering detection, understanding, and mitigation strategies.

AI/ML arXiv cs.AI

Trust-free Personalized Decentralized Learning

TPFed is introduced as a trust-free, decentralized federated learning framework using blockchain and LSH to enable personalized collaboration without a central aggregator.

AI/ML arXiv cs.AI

Leveraging Metamemory Agent for Enhanced Data-Free Code Generation in Large Language Models

A new metamemory agent is proposed to improve LLM code generation in data-free scenarios by guiding the model to recall and evaluate its own internal knowledge.

AI/ML arXiv cs.AI

The relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer

An analysis of OpenAI's o3-mini model shows it achieves higher accuracy with shorter reasoning chains compared to o1-mini, suggesting more efficient use of test-time compute.

AI/ML arXiv cs.AI

MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning

Introduces MLLM-LLaVA-FL, a federated learning framework that uses multimodal large language models at the server side to handle data heterogeneity and long-tail distributions.

AI/ML arXiv cs.AI

Platonic Representations for Poverty Mapping: Unified Vision-Language Codes or Agent-Induced Novelty?

Develops a multimodal framework for poverty mapping using satellite imagery and LLM-generated text, releasing a dataset of 60,000 DHS clusters.

AI/ML arXiv cs.AI

Base Models Know How to Reason, Thinking Models Learn When

Analyzes the difference between base and 'thinking' models, finding that RL primarily teaches orchestration heuristics for existing mechanisms while SFT installs new ones.

AI/ML arXiv cs.AI

Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents

Introduces PROBE, a benchmark to measure proactive problem solving in LLM agents, revealing that even state-of-the-art models struggle with autonomous resolution.

AI/ML arXiv cs.AI

When Assisting One Disempowers Another

Formalizes 'bystander disempowerment,' where an AI assistant optimizing for one user inadvertently reduces the agency of nearby bystanders.

AI/ML arXiv cs.AI

VASP Agent: An Agentic Framework for Autonomous First-principles Calculations

Presents VASP Agent, an agentic framework that automates first-principles materials calculations by combining domain skills with deterministic tools.

Software Engineering arXiv cs.AI

Implementing Metric Temporal Answer Set Programming

Proposes a computational approach to Metric Answer Set Programming (ASP) that uses difference constraints to decouple temporal constraints from time precision.

Tech Business/VC arXiv cs.AI

Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique

Describes an agentic AI system for insurance underwriting using adversarial self-critique to reduce hallucinations and maintain human-in-the-loop authority.

AI/ML arXiv cs.AI

Self-Routing: Parameter-Free Expert Routing from Hidden States

Introduces Self-Routing, a parameter-free MoE routing mechanism that uses token hidden states directly as logits, eliminating the need for a learned router.

AI/ML arXiv cs.AI

Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents

Investigates the 'faithfulness gap' in LLM agents using a poker simulator, finding that agents often fail in the reasoning-to-conclusion step despite reliable conclusion-to-action execution.

AI/ML Hacker News

I Think I Have LLM Burnout

A community discussion on Hacker News regarding the feeling of burnout associated with the rapid pace and hype of Large Language Models.

Cybersecurity Hacker News

Apache Shiro security framework releases 3.0.0

Apache Shiro, a Java security framework, has released version 3.0.0.

Tech Business/VC TechCrunch

Truecaller clashes with India’s telecom regulator over anti-spam rules

Truecaller is in a dispute with India's telecom regulator over the effectiveness of anti-spam rules and business number series.

AI/ML arXiv cs.AI

Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

Introduction of DataGovBench, a benchmark using governmental open data to evaluate LLMs' ability to handle real-world data complexities and exploratory analysis.

AI/ML arXiv cs.AI

AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models

Research introducing AirflowAttack, a novel adversarial attack targeting infrared remote-sensing vision-language models by simulating thermal-airflow turbulence.