Software Engineering Hacker News

Logical Physical Clocks and Consistent Snapshots in Globally Distributed DB [pdf]

A technical discussion or paper regarding the implementation of logical physical clocks and consistent snapshots for globally distributed databases.

AI/ML arXiv cs.AI

The CRISTAL Method: Neurosymbolic analysis from AI-synthesized world models

Introduces the CRISTAL Method, a neurosymbolic framework that combines LLMs with probabilistic programs for reproducible and justified investment analysis.

AI/ML arXiv cs.AI

Beyond Triplet Plausibility: Relation Set Completion in Knowledge Graphs

Proposes RelSetE, a Relation Set Embedding model designed to improve knowledge graph completion by reasoning about missing relations compatible with an entity.

AI/ML arXiv cs.AI

AI Training Manager: Bounded Closed-Loop Control of Adaptive Training Recipes

Presents the AI Training Manager, an LLM-based supervisory controller that adaptively adjusts training parameters like learning rate and regularization in real-time.

AI/ML arXiv cs.AI

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing

Introduces SafePyramid, a hierarchical benchmark for evaluating in-context policy guardrailing in LLMs, revealing significant struggles in complex rule reasoning.

AI/ML arXiv cs.AI

A causal modeling perspective on decision theory

Proposes a formal framework for decision theory using nonparametric structural equation models to create a unified modeling language for agents and counterfactuals.

AI/ML arXiv cs.AI

HippoSpark: An On-Demand Experience System for LLM Reasoning

Introduces HippoSpark, an on-demand state-level experience retrieval system that enhances LLM reasoning by providing actionable guidance at critical bottlenecks.

AI/ML arXiv cs.AI

SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon CivRealm Strategy Planning

Presents SAGA, a multi-agent framework using scene graphs and a dual-horizon feedback loop to improve long-horizon strategic planning in complex games like FreeCiv.

AI/ML arXiv cs.AI

First-Order Temporal Logic Tensor Networks

Introduces First-Order Temporal Logic Tensor Networks (FOT-LTN), extending Logic Tensor Networks to handle objects whose properties change over time.

AI/ML arXiv cs.AI

Exploration and Online Transfer with Behavioral Foundation Models

Explores using Behavioral Foundation Models to generate exploration policies for online transfer in zero-shot reinforcement learning via a bandit-like approach.

AI/ML arXiv cs.AI

Budgeted Act-or-Defer Multi-Agent LLM Deliberation with Local Reliability Bounds

Proposes a budgeted act-or-defer framework for multi-agent LLM deliberation that uses local reliability bounds to decide when to act or escalate to human review.

AI/ML arXiv cs.AI

Safety from Honesty in a Disinterested AI Predictor

Presents a formal safety argument for the Scientist AI (SAI) Predictor, aimed at preventing implicit agency and ensuring the model remains a disinterested predictor.

AI/ML arXiv cs.AI

Diversity is the Strength of the AI Crowd

Analyzes AI forecasting ensembles and finds that diversity among models, specifically combining complementary errors, is more critical for accuracy than simply increasing sample size.

AI/ML arXiv cs.AI

Sample-Efficient Learning of Probabilistic Causes for Reachability in Markov Decision Processes with Probabilistic Guarantees

Introduces a learning approach with probabilistic guarantees for identifying 'probability-raising' causes in unknown Markov Decision Processes (MDPs).

AI/ML arXiv cs.AI

Toward Secure and Reliable PDDL Formalization of Large Language Models with Planner-in-the-Loop Feedback

Presents NL-PDDL-Bench, a benchmark for NL-to-PDDL specification, and a planner-in-the-loop framework to improve the reliability of LLM-based planning.

AI/ML arXiv cs.AI

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots

Introduces GUICrafter, a weakly-supervised GUI agent that uses massive unannotated screenshots and curriculum learning to reduce reliance on human labels.

AI/ML arXiv cs.AI

DeepTrans Studio: Turning Expert Interventions into Shared Team Knowledge in Agentic Translation Workflows

Describes DeepTrans Studio, a collaborative workspace for agentic translation that turns expert human corrections into shared team knowledge.

Open Source arXiv cs.AI

DEEPMED Search: An Open-Source Agentic Platform for Medical Deep Research with Introspective Verification

Presents DEEPMED Search, an open-source agentic platform for medical research that uses an introspective verification module to validate retrieved evidence.

Cybersecurity arXiv cs.AI

Rethinking Generative Reconstruction Attacks against Graph Neural Network Models

Explores vulnerabilities in Graph Neural Networks (GNNs) by introducing two novel generative reconstruction attacks (GLC and ELC) to leak sensitive training data.

AI/ML arXiv cs.AI

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

Introduces CLQT, a closed-loop benchmark for evaluating LLM portfolio-management agents based on strategy consistency and capability diagnosis rather than simple returns.