AI/ML arXiv cs.AI

RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories

Introduction of RippleBench, a pipeline for analyzing the side-effects of model unlearning and editing in LLMs.

AI/ML arXiv cs.AI

Towards Understanding What State Space Models Learn About Code

A systematic analysis of State Space Models (SSMs) in code understanding, introducing SSM-Interpret to analyze spectral shifts during fine-tuning.

Other Hacker News

Update on Ocean Observatories Initiative

An update on the status and operations of the Ocean Observatories Initiative.

Other Hacker News

It doesn't matter if it works

A discussion regarding the philosophical or technical implications of whether a system's internal workings matter if the end result is successful.

AI/ML arXiv cs.AI

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

Introduces ScenA, a text-to-audio model that generates multi-speaker audio scenes with natural ambient noise and overlapping speech using flow-matching.

AI/ML arXiv cs.AI

UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning

Presents UBP2, a model-based reinforcement learning approach that uses uncertainty-balanced preference planning to improve sample efficiency in preference-based RL.

AI/ML arXiv cs.AI

Large-Scale OD Matrix Estimation with A Deep Learning Method

Proposes a hybrid deep learning and numerical optimization method for more accurate and real-time origin-destination (OD) matrix estimation in transport systems.

AI/ML arXiv cs.AI

Recursive Joint Simulation in Games

Explores the use of recursive joint simulation between AI agents to achieve more cooperative outcomes in strategic game-theoretic settings.

AI/ML arXiv cs.AI

Fully Geometric Multi-Hop Reasoning on Knowledge Graphs with Transitive Relations

Introduces GeometrE, a geometric embedding method for multi-hop reasoning on knowledge graphs that maps logical operations to pure geometric transformations.

AI/ML arXiv cs.AI

PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation

Presents PosterForest, a training-free framework that uses a hierarchical multi-agent collaboration to automate scientific poster generation.

AI/ML arXiv cs.AI

Structured Cognitive Loop for Behavioral Intelligence in Large Language Model Agents (Extended Revision: From Behavioral Architecture to Epistemic Accountability)

Proposes the Structured Cognitive Loop (SCL) architecture to improve accountability and success rates in LLM agents by separating cognition, memory, and control.

AI/ML arXiv cs.AI

The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs

Analyzes how long-term user memory in LLMs can create the 'personalization trap,' where demographic profiles bias emotional reasoning and reinforce social inequalities.

Other Hacker News

The AirPods Effect

A discussion on the 'AirPods Effect', likely referring to how a specific product's success changes consumer behavior or industry standards.

Software Engineering Hacker News

Flip TABLE: storing arbitrary data in iNaturalist

A technical discussion about storing arbitrary data within the iNaturalist platform using a 'Flip TABLE' approach.

Other Ars Technica

FDA advisors unanimously vote to approve Moderna's mRNA after agency drama

FDA advisors have unanimously voted to approve Moderna's mRNA vaccine following previous administrative delays.

Hardware/Chips Ars Technica

As China looms, Taiwan makes more drones for defense and the US military

Taiwan is increasing drone production for national defense and for export to the US military amid rising tensions with China.

AI/ML VentureBeat

Anthropic's Claude Code Artifacts update brings live, shared dashboards and interactive workspaces to enterprises

Anthropic introduces 'Artifacts' for Claude Code, enabling the creation of live, shared, and interactive HTML workspaces for enterprise teams.

AI/ML arXiv cs.AI

A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2

Researchers introduce a multi-domain benchmark to detect AI-generated text-rich images (like receipts and infographics) from GPT-Image-2.

AI/ML arXiv cs.AI

Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA

An empirical study on adapting LLMs for French medical QA, comparing continual pretraining (CPT) and supervised fine-tuning (SFT).

AI/ML arXiv cs.AI

Correct Yourself, Keep My Trust: How Self-Correction and Social Connection Shape Credibility in Social Chatbots

Research on social chatbots suggests that self-correction is more effective at maintaining user trust than external corrections.