AI/ML arXiv cs.AI

Animation, Verification and Visualisation of Prolog Transition Systems with ProB

Researchers present updates to ProB, a Prolog-based model checker and animator, featuring improved state visualization and statistical simulation.

Software Engineering arXiv cs.AI

Chess\_db: A framework for working with large chess game datasets

Researchers introduce Chess_db, a logic programming suite for managing and querying large chess game datasets using key-value databases.

AI/ML arXiv cs.AI

Case study: solving P-99 with LPTP and an LLM

A study demonstrates using Claude LLM to generate, test, and formally prove Prolog code for the Ninety-Nine Prolog Problems (P-99) set.

AI/ML arXiv cs.AI

Declarative Problem Solving in UAM Strategic Deconfliction

A proposal uses Answer Set Programming (ASP) for strategic deconfliction in Urban Air Mobility to manage drone and air taxi traffic.

AI/ML arXiv cs.AI

Towards a Certifying Grounder

Researchers introduce CertiFOX, a framework for certifying the grounding process in first-order logic model expansion.

AI/ML arXiv cs.AI

Hybrid MKNF with Classical Negation in the Rule Component

A research paper explores extending Hybrid MKNF knowledge bases to support classical negation within the rule component.

AI/ML arXiv cs.AI

Explainability Framework for Policy-Aware Autonomous Agents

A framework is proposed to provide explainable behavior for autonomous agents using Answer Set Programming and Python.

AI/ML arXiv cs.AI

Explainable Belief Harmonization under Dynamic Epistemic Partitions

A formal framework is presented for managing multi-agent belief harmonization under dynamic changes in observational capacity.

Software Engineering Hacker News

How to Write a Quine

A discussion on the concept and implementation of quines, which are computer programs that take no input and produce a copy of their own source code as their only output.

AI/ML arXiv cs.AI

GlucoTune: A Unified Framework for Blood Glucose Preprocessing, Forecasting, and Benchmarking in Diabetes

Introduction of GlucoTune, an extensible framework for preprocessing, forecasting, and benchmarking blood glucose time-series data for diabetes management.

AI/ML arXiv cs.AI

Relative Value Learning

Proposed Relative Value Learning (RV) framework for reinforcement learning that focuses on learning value differences rather than absolute state values, achieving competitive results on Atari benchmarks.

Hardware/Chips arXiv cs.AI

Hardware-Software Co-Design for Float16 On-Device Training on RISC-V Single-Core

An open-source framework for on-device training on RISC-V single-core using Float16 to reduce memory footprint by 50% with minimal performance loss.

AI/ML arXiv cs.AI

Demographically-Informed Heat-Mortality Risk Curves via Risk Graph Neural Networks

Development of Risk Graph Neural Networks (RGNNs) to improve heat-mortality risk estimation by integrating demographic and geographic context.

AI/ML arXiv cs.AI

One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies

Introduction of RegretBench, a multi-turn benchmark evaluating the clarification policies of LLMs based on a regret-based objective.

AI/ML arXiv cs.AI

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

CRAG-MM-Diagnostics is a diagnostic benchmark for Knowledge-Intensive VQA that isolates failures in grounding, identification, and retrieval stages.

AI/ML arXiv cs.AI

Representative Sets in Propositional Abduction

Theoretical research on propositional abduction, exploring representative sets of explanations and establishing links between non-monotonic reasoning and coding theory.

Software Engineering arXiv cs.AI

Case study: proving sqrt(2) irrational with LPTP and an LLM

A case study using the Logic Program Theorem Prover (LPTP) and an LLM to generate a formal proof for the irrationality of sqrt(2).

Software Engineering arXiv cs.AI

Encoding Event-B Proof Rules in Prolog: An Interactive Sequent Prover for ProB

An interactive sequent prover for Event-B developed in Prolog, integrated into ProB to provide proof tree visualization and better educational control.

AI/ML Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

Opus 5 has reached the top position on the Artificial Analysis Intelligence Leaderboard, indicating high performance in AI model benchmarking.

Tech Business/VC TechCrunch

Prentis, new AI lab co-founded by Reid Hoffman, Marc Pincus in talks to raise $100M

Reid Hoffman and Marc Pincus are launching Prentis, a new AI lab focused on automating routine computer tasks beyond simple coding.