AI/ML arXiv cs.AI

RoME: Robust Mixture of Low-Rank Experts against Multiple Adversarial Perturbations

RoME introduces a Robust Mixture of Low-Rank Experts to improve AI model robustness against multiple adversarial perturbations.

AI/ML arXiv cs.AI

LLM-Guided Measurement Credibility Correction for Trustworthy Industrial Process Inference

LLM-Guided Measurement Credibility Correction (MCC) uses LLMs to interpret process documents and correct industrial sensor data for better prediction accuracy.

AI/ML arXiv cs.AI

x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability

Truncated Jump Sampling (TJS) enables faster generation in diffusion and flow matching models by predicting the endpoint without retraining.

Software Engineering arXiv cs.AI

Static Metrics Are Insufficient: Predicting Java Method Energy Usage with Execution Time

Research indicates that combining static source code metrics with execution time significantly improves the prediction of Java method energy usage.

AI/ML arXiv cs.AI

Evaluating Fine-Tuning and Metrics for Neural Decompilation of Dart AOT Binaries

A study on neural decompilation of Dart AOT binaries finds that pass@k is the most valid metric and introduces the HumanEval-Dart benchmark.

AI/ML arXiv cs.AI

Self-Supervised Implicit CEST Reconstruction via Physics-Informed Lorentz Encoding

Lorentz Encoding (LE) is a physics-informed framework for self-supervised reconstruction of CEST MRI spectra to reduce scan times.

Software Engineering arXiv cs.AI

Property-Driven Synthetic Data Engineering for Data-Scarce Software Systems: Reflections from the Breast Cancer Domain

The authors propose 'property-driven synthetic data engineering' to address data scarcity in sensitive domains like breast cancer radiotherapy software.

Other Hacker News

My road trip with the do-gooding cactus smugglers

A personal narrative about a road trip with cactus smugglers, shared on Hacker News.

AI/ML arXiv cs.AI

InfluMatch: Frontier-Quality KOL Search at 4B-Model Cost

Introduces InfluMatch, a cost-effective three-stage cascade system for influencer search using small open-weight models instead of expensive frontier LLMs.

AI/ML arXiv cs.AI

Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search

Studies the trade-off between retrieval effectiveness and faithfulness when using LLM-generated synthetic metadata for RDF dataset search.

AI/ML arXiv cs.AI

MCP-Enabled Agentic AI for Autonomous IPoDWDM Network Lifecycle Automation

A demo of an MCP-enabled agentic AI architecture for autonomous control and lifecycle automation of vendor-agnostic IPoDWDM networks.

AI/ML arXiv cs.AI

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention

Proposes Multi-Token Localized Attention (MTLA), a training-free method to reduce hallucinations in multimodal LLMs by measuring attention strength to claimed regions.

AI/ML arXiv cs.AI

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

Introduces PluraMath, an open-source extension of the PolyMath dataset to 18 underrepresented languages to evaluate mathematical reasoning in LLMs.

Software Engineering arXiv cs.AI

Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development

Evaluates Prompt Coach, an agentic tutor that provides Socratic guidance within the IDE to help software developers learn prompt engineering.

AI/ML arXiv cs.AI

From Blueprint to Reality: Modeling and Applying Putnam's Social Capital Theory with LLM-based Multi-agent Simulations

Introduces SocaSim, an LLM-based multi-agent simulation framework for studying social capital theory and collective action.

AI/ML arXiv cs.AI

PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet

Proposes PVCap, a 3D dense captioning method utilizing PseudoCap for data augmentation and VoxelCapNet for improved semantic extraction.

Software Engineering arXiv cs.AI

Agents That Teach: Towards Designing Incidental Learning Back into AI-Assisted Software Development

Discusses the 'Knowledge Debt' caused by AI coding agents and proposes SHIELD, a system designed to integrate incidental learning back into AI-assisted development.

Software Engineering Hacker News

Rewriting Bun in Rust

A discussion regarding the possibility or process of rewriting the Bun runtime in Rust.

Other Hacker News

New Sweden: the US's long-lost 'secret' colony

An article discussing the history of 'New Sweden', a former Swedish colony in the United States.

Tech Business/VC TechCrunch

Despite ‘misgivings,’ judge approves Elon Musk’s $1.5 million SEC settlement

A judge has approved Elon Musk's $1.5 million settlement with the SEC regarding his Twitter stake disclosures.