AI/ML arXiv cs.AI

Disentangling Aleatoric and Epistemic Uncertainty in Physics-Informed Neural Networks. Application to Insulation Material Degradation Prognostics

Introduces a Bayesian Physics-Informed Neural Network (B-PINN) framework to jointly model aleatoric and epistemic uncertainty for insulation material degradation prognostics.

Software Engineering arXiv cs.AI

The $\mathbf{P}$-Completeness of Inverted Index Traversal: On the Complexity of Evaluating Boolean Query DAGs

Proves the P-completeness of inverted index traversal for Boolean query DAGs and introduces ComputePN, a sparsity-aware evaluation algorithm to make such queries tractable.

AI/ML arXiv cs.AI

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

Analyzes the 'narcissism' of LLM evaluators, finding that self-preference may often be a result of evaluator quality rather than a systematic bias toward their own outputs.

AI/ML arXiv cs.AI

Toward Autonomous O-RAN: A Multi-Scale Agentic AI Framework for Real-Time Network Control and Management

Proposes a multi-scale agentic AI framework for O-RAN, using a hierarchy of LLM, SLM, and Physical-layer Foundation Model agents for real-time network control.

Software Engineering Hacker News

Show HN: Brain Frog – Can you be random enough for 11 lines of JavaScript?

A small JavaScript-based game or challenge called Brain Frog that tests the user's ability to be random.

AI/ML arXiv cs.AI

SEAL: Searching Expandable Architectures for Incremental Learning

Introduction of SEAL, a NAS-based framework for data-incremental learning that adapts model structure dynamically to balance plasticity and stability.

AI/ML arXiv cs.AI

Graph Alignment for Benchmarking Graph Neural Networks and Learning Positional Encodings

A new benchmarking methodology for Graph Neural Networks based on the graph alignment problem, including an open-source Python package.

AI/ML arXiv cs.AI

Render-FM: Feedforward Model for Real-time Photorealistic Volumetric Rendering

Render-FM is a feedforward model that enables real-time photorealistic volumetric rendering of CT scans, providing a 500x speedup over traditional optimization.

AI/ML arXiv cs.AI

Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training

BBoxER is an evolutionary black-box optimization method for LLM post-training designed to improve privacy and robustness against data poisoning.

AI/ML arXiv cs.AI

FISHER: A Foundation Model for Multi-Modal Industrial Signal Comprehensive Representation

FISHER is a foundation model for multi-modal industrial signal representation that outperforms larger encoders while maintaining diagnostic accuracy.

AI/ML arXiv cs.AI

Rule2Text: A Framework for Generating and Evaluating Natural Language Explanations of Knowledge Graph Rules

Rule2Text is a framework that uses LLMs to convert complex knowledge graph rules into natural language explanations.

Cybersecurity arXiv cs.AI

FALCON: Transforming Cyber Threat Intelligence into Deployable IDS Rules with Self-Reflection

FALCON is an agentic framework that automates the transformation of Cyber Threat Intelligence into deployable IDS rules for Snort and YARA.

AI/ML arXiv cs.AI

Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

Research on using lightweight steering vectors to mitigate self-preference bias in LLM evaluators at inference time.

AI/ML arXiv cs.AI

VoltanaLLM: Energy-Efficient and SLO-Aware Disaggregated LLM Serving via Adaptive Frequency Control and State-Space Routing

VoltanaLLM is a system for energy-efficient LLM serving that uses adaptive frequency control and state-space routing to reduce energy consumption.

Other Hacker News

GTA 6 Physical Copies Won't Include a Disc, Will Just Be a Code in a Box

Reports on GTA 6 physical editions lacking actual discs, providing only digital download codes instead.

Other Hacker News

Medical students are using popular research tool to pump out misleading studies

Discusses how medical students are leveraging popular research tools to generate misleading academic studies.

AI/ML arXiv cs.AI

Impatient Bandits: Optimizing for the Long-Term Without Delay

Introduces a bandit algorithm and predictive model to optimize for long-term user satisfaction in recommender systems despite delayed rewards.

AI/ML arXiv cs.AI

Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions

Proposes RV-Bench, a benchmark using random variable questions to evaluate the genuine mathematical reasoning capabilities of LLMs and detect data contamination.

AI/ML arXiv cs.AI

Societal Alignment Frameworks Can Improve LLM Alignment

Argues for incorporating societal alignment frameworks (social, economic, contractual) to improve the alignment of large language models.

AI/ML arXiv cs.AI

Reward-Centered ReST-MCTS: A Robust Decision-Making Framework for Robotic Manipulation in High Uncertainty Environments

Presents Reward-Centered ReST-MCTS, a framework to improve robotic manipulation decision-making in high-uncertainty environments using a test-time verifier.