Software Engineering arXiv cs.AI

Pramana: A Composable, Domain-Specific Backend for Empirical Networking Research

Pramana is introduced as a composable backend designed to accelerate empirical networking research by bridging the gap between hypothesis and data generation.

AI/ML arXiv cs.AI

High-Order Markov Blanket Discovery via a k-Order Relaxation of the Faithfulness Assumption

The k-order Markov blanket (kOMB) algorithm is proposed to discover graphical Markov blankets while relaxing the faithfulness assumption to handle parity-type relations.

Other Hacker News

"the very foundation of modern academia has been blown to bits"

A Hacker News discussion regarding the perceived collapse of the foundations of modern academia due to external pressures or AI.

AI/ML arXiv cs.AI

Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models

Research showing that 'blind resampling' (retrying without looking at previous failures) is more efficient and often more effective than self-repair in small code LLMs.

AI/ML arXiv cs.AI

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

A proposed systems framework and graded trustworthiness levels for embodied AI to ensure safe and reliable physical interaction.

AI/ML arXiv cs.AI

A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment

A hybrid deep learning approach using Mask R-CNN and color-based segmentation to quantify skin exposure for safety assessments.

AI/ML arXiv cs.AI

Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition

An analysis of the structural and cognitive convergence between Large Language Models and human cognitive organization.

Cybersecurity arXiv cs.AI

(EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigations

Introduction of (EC)2, a multi-agent LLM framework for providing event-centric, verifiable explanations for cybersecurity alerts.

AI/ML arXiv cs.AI

Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges

A systematic survey and taxonomy of Multi-Agent Debate (MAD) strategies to improve LLM robustness and accuracy.

Software Engineering arXiv cs.AI

Model-Driven Requirements Configuration with Three-Valued Uncertainty Scoring

A neuro-symbolic multi-agent architecture that combines LLMs with a deterministic symbolic validator to ensure structural integrity in requirements engineering.

AI/ML arXiv cs.AI

Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech

Study on generating personalized, contextualized counterspeech to mitigate online toxicity more effectively than generic responses.

AI/ML arXiv cs.AI

Top-$k$ Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection

Introduction of THV-UCB, an optimistic algorithm for multi-objective slate selection to approximate Pareto frontiers in bandit problems.

Hardware/Chips Hacker News

Show HN: Cubic Doggo 06R: 12-DOF 4-Legged Robot with IMU

A demonstration of Cubic Doggo 06R, a 12-degree-of-freedom 4-legged robot equipped with an IMU.

AI/ML arXiv cs.AI

Optimizing Sensor Placement for Hydrogen Leak Detection in Enclosed Infrastructure: A Comparative Study Using CFD-informed Genetic Algorithm and DeepSets Neural Surrogate

A study proposing a computational framework using CFD, genetic algorithms, and DeepSets neural surrogates to optimize sensor placement for hydrogen leak detection.

AI/ML arXiv cs.AI

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models

Introduction of the Reasoning Answer Faithfulness Score (RAFS), a reference-free metric to detect silent reasoning failures in LLM mathematical chain-of-thought evaluations.

AI/ML arXiv cs.AI

Weight and Height Estimation from a Single Human Image Captured in the Wild

Research on estimating human weight and height from single images using deep neural networks and a newly proposed image dataset.

AI/ML arXiv cs.AI

TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions

TraceCLIP is a training-free framework that recovers local semantic evidence from CLIP's global representations to improve zero-shot semantic segmentation.

AI/ML arXiv cs.AI

GPT-Red: Automated Red Teaming via Self-Play at Scale

GPT-Red is an automated red-teaming agent trained via self-play to discover novel prompt injection attacks and improve LLM robustness.

Open Source Hacker News

Show HN: Gander, an Android file viewer that asks for no permissions at all

Gander is an Android file viewer application designed to operate without requiring any system permissions.

AI/ML arXiv cs.AI

Do Methods Support the Claims? Intra-Paper Verification for Peer Review

Researchers introduce a framework for intra-paper claim verification to help LLMs identify when a scientific paper's claims are not supported by its own methods.