AI/ML arXiv cs.AI

Diverse Evidence, Better Forecasts: Multi-Agent Deliberation Under Information Asymmetry

Introduction of InfoDelphi, a framework that uses designed information asymmetry among multiple LLM agents to improve forecasting accuracy and reduce herding.

AI/ML arXiv cs.AI

Separating Expert Retention from Autonomous Source Inference in Raw-ECG-Replay-Free Continual ECG Deployment

Research on an incremental expert bank for ECG deployment that separates expert retention from autonomous source inference to handle new data sources without replaying raw data.

AI/ML arXiv cs.AI

Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing

Introduction of Goggles, a pretrained module that edits finetuning gradients to impose an epistemic frame, preventing models from believing fictional data during training.

AI/ML arXiv cs.AI

COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows

Presentation of COMFYCLAW, an agentic skill evolution harness for ComfyUI that uses typed graph editing and VLM verification to build reusable image generation skills.

AI/ML arXiv cs.AI

Generic Expert Coverage for Pruning SparseMixture-of-Experts Language Models

A new expert pruning method called Generic TB-Coverage that preserves cross-corpus utility to reduce redundancy in Sparse Mixture-of-Experts (MoE) models without downstream data.

AI/ML arXiv cs.AI

Distributionally Robust Listwise Preference Optimization

Proposed a distributionally robust listwise preference optimization objective to improve LLM alignment under ranking-label uncertainty and noise.

Cybersecurity arXiv cs.AI

DRL-CLBA: A Clean Label Backdoor Attack for Speech Classification via DDPG Reinforcement Learning

Development of DRL-CLBA, a clean label backdoor attack for speech classification using DDPG reinforcement learning and deep audio steganography.

Software Engineering arXiv cs.AI

Reformalization of the Jordan Curve Theorem

A case study on the reformalization of the Jordan Curve Theorem across different proof assistants including Mizar, Lean, HOL Light, and Agda.

AI/ML arXiv cs.AI

Meta-Benchmarks for Financial-Services LLM Evaluation

A meta-benchmarking framework for financial-services LLMs that organizes hundreds of benchmarks into banking business domains using a multiplicative weighting scheme.

AI/ML Hacker News

14× faster embeddings: how we rebuilt the ONNX path in Manticore

Manticore details how they achieved a 14x speedup in embeddings by rebuilding their ONNX execution path.

Other Hacker News

Cowboys, Frontiersmen, Settlers, Townspeople, Cityfolk

A discussion on the archetypes of pioneers and city dwellers, likely exploring sociological or historical transitions.

Cybersecurity TechCrunch

Politician who investigated spyware abuses had his phone hacked with Pegasus spyware

A European politician investigating spyware was targeted by Pegasus spyware, highlighting ongoing abuses by NSO Group customers.

AI/ML arXiv cs.AI

Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation

Introduces Hawk, a training-free framework for generating high-performance NPU kernels by incorporating hardware-aware knowledge.

AI/ML arXiv cs.AI

Safe and Adaptive Cloud Healing: Verifying LLM-Generated Recovery Plans with a Neural-Symbolic World Model

Presents PASE, a neuro-symbolic framework that uses LLMs and a world model to generate and verify autonomous recovery plans for cloud faults.

AI/ML arXiv cs.AI

SemHash-LLM: A Multi-Granularity Semantic Hashing Framework for Document Deduplication

Proposes SemHash-LLM, a multi-granularity semantic hashing framework designed for efficient large-scale document deduplication.

AI/ML arXiv cs.AI

Profit-Based Counterfactual Explanations for Product Improvement: A Case Study of Manga Sales in Japan

Introduces PBCE, a framework for profit-based counterfactual explanations to optimize product attributes based on profit maximization.

AI/ML arXiv cs.AI

Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling

Introduces C3RL for better LLM confidence calibration and CAS for adaptive test-time scaling to reduce inference costs.

AI/ML arXiv cs.AI

Spatial Support Matters: Geometry-Aware Graph Fusion for Rainfall Field Reconstruction

Proposes a geometry-aware graph neural network for rainfall field reconstruction using heterogeneous spatial supports.

AI/ML arXiv cs.AI

Autonomous discovery of traffic laws with AI traffic scientists

Presents TrafficSci, an agentic AI system capable of autonomously discovering universal traffic laws through an iterative scientific workflow.

Other Hacker News

The Free Market Lie: Why Switzerland Has 25 Gbit Internet and America Doesn't

An analysis of the differences in internet infrastructure and accessibility between Switzerland and the United States, challenging free-market assumptions.