All Articles
17528 articles total
A Three-Layer Framework for AI in Scientific Discovery
A proposed three-layer framework for AI in scientific discovery, emphasizing 'model formation' via qualitative reasoning over simple search or execution.
Replication in Visual Diffusion Models: A Survey and Outlook
A comprehensive survey on the phenomenon of replication (memorization) in visual diffusion models, covering detection, understanding, and mitigation strategies.
Trust-free Personalized Decentralized Learning
TPFed is introduced as a trust-free, decentralized federated learning framework using blockchain and LSH to enable personalized collaboration without a central aggregator.
Leveraging Metamemory Agent for Enhanced Data-Free Code Generation in Large Language Models
A new metamemory agent is proposed to improve LLM code generation in data-free scenarios by guiding the model to recall and evaluate its own internal knowledge.
The relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer
An analysis of OpenAI's o3-mini model shows it achieves higher accuracy with shorter reasoning chains compared to o1-mini, suggesting more efficient use of test-time compute.
MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning
Introduces MLLM-LLaVA-FL, a federated learning framework that uses multimodal large language models at the server side to handle data heterogeneity and long-tail distributions.
Platonic Representations for Poverty Mapping: Unified Vision-Language Codes or Agent-Induced Novelty?
Develops a multimodal framework for poverty mapping using satellite imagery and LLM-generated text, releasing a dataset of 60,000 DHS clusters.
Base Models Know How to Reason, Thinking Models Learn When
Analyzes the difference between base and 'thinking' models, finding that RL primarily teaches orchestration heuristics for existing mechanisms while SFT installs new ones.
Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
Introduces PROBE, a benchmark to measure proactive problem solving in LLM agents, revealing that even state-of-the-art models struggle with autonomous resolution.
When Assisting One Disempowers Another
Formalizes 'bystander disempowerment,' where an AI assistant optimizing for one user inadvertently reduces the agency of nearby bystanders.
VASP Agent: An Agentic Framework for Autonomous First-principles Calculations
Presents VASP Agent, an agentic framework that automates first-principles materials calculations by combining domain skills with deterministic tools.
Implementing Metric Temporal Answer Set Programming
Proposes a computational approach to Metric Answer Set Programming (ASP) that uses difference constraints to decouple temporal constraints from time precision.
Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique
Describes an agentic AI system for insurance underwriting using adversarial self-critique to reduce hallucinations and maintain human-in-the-loop authority.
Self-Routing: Parameter-Free Expert Routing from Hidden States
Introduces Self-Routing, a parameter-free MoE routing mechanism that uses token hidden states directly as logits, eliminating the need for a learned router.
Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents
Investigates the 'faithfulness gap' in LLM agents using a poker simulator, finding that agents often fail in the reasoning-to-conclusion step despite reliable conclusion-to-action execution.
I Think I Have LLM Burnout
A community discussion on Hacker News regarding the feeling of burnout associated with the rapid pace and hype of Large Language Models.
Apache Shiro security framework releases 3.0.0
Apache Shiro, a Java security framework, has released version 3.0.0.
Truecaller clashes with India’s telecom regulator over anti-spam rules
Truecaller is in a dispute with India's telecom regulator over the effectiveness of anti-spam rules and business number series.
Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities
Introduction of DataGovBench, a benchmark using governmental open data to evaluate LLMs' ability to handle real-world data complexities and exploratory analysis.
AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models
Research introducing AirflowAttack, a novel adversarial attack targeting infrared remote-sensing vision-language models by simulating thermal-airflow turbulence.