AI/ML Hacker News

DSLs Enable Reliable Use of LLMs

A technical exploration of how Domain Specific Languages (DSLs) can be used to ensure reliable and structured interaction with Large Language Models.

AI/ML Hacker News

Societal Impacts: Claude's values across models and languages

A research study on the societal impacts of Claude AI, examining how the model's values shift across different languages and models.

Other Hacker News

Telegram Serverless

A discussion or article about the architecture or serverless-style implementation of Telegram services.

Other Hacker News

Floating Companion: Exploring Design Space for Soft Floating Robots in Indoor

Discussion on the design space for soft floating robots intended for use in indoor environments.

Hardware/Chips Hacker News

Show HN: Web App Uses RTL-SDR to Align HDTV Antenna

A web application that utilizes RTL-SDR hardware to help users align their HDTV antennas.

AI/ML arXiv cs.AI

Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs

EG-VAR is a Lean 4-based architecture that eliminates LLM hallucinations in empirical inference by using a kernel to verify tool-attested claims.

AI/ML arXiv cs.AI

Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference

Jetson-PI improves VLA model deployment on Jetson Orin through foresight-aligned asynchronous inference to reduce latency and increase control frequency.

AI/ML arXiv cs.AI

Text-Aided Multi-Modal Panoptic Symbol Spotting for CAD Floor Plan Drawings

TextCAD is a multimodal framework that combines graphical primitives and textual annotations to improve symbol spotting in CAD floor plans.

AI/ML arXiv cs.AI

From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation

CARE-PPO is an RL framework that uses a confidence-aligned reward to improve quantitative prediction and confidence estimation in LLMs.

AI/ML arXiv cs.AI

Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts

EcoSpec is a cost-aware speculative decoding framework for MoE models that reduces memory traffic and improves decoding speed by optimizing expert activation.

Software Engineering arXiv cs.AI

Line-Anchored Feedback Cuts Token Costs and Improves Correctness in AI Code Editing

Line-anchored feedback via the FileMark extension reduces token costs and increases correctness in AI-driven code editing compared to holistic prompts.

Cybersecurity arXiv cs.AI

Bulkhead: Automated Semantic Detection and Remediation of Container Escape Vulnerabilities

Bulkhead is an automated framework using LLMs and formal methods to detect and remediate container escape vulnerabilities caused by path traversal.

AI/ML arXiv cs.AI

Learning-based Probabilistic Load Forecasting with Post-hoc and In-model Uncertainty

Research on probabilistic load forecasting for smart buildings, comparing post-hoc and in-model uncertainty estimation with various DL backbones.

Tech Business/VC TechCrunch

Backed by $60M in funding, Oak steps out of stealth to fix the identity mess that AI agents are making worse

Identity management startup Oak emerges from stealth with $60 million in funding to address identity challenges created by AI agents.

Hardware/Chips Ars Technica

How hard is it to build orbital data centers, actually?

An exploration of the technical challenges and costs associated with building orbital data centers, focusing on cooling and radiators.

Other Ars Technica

Sotheby's big T. rex auction raises concerns hype and wealth are upending science

High-wealth private buyers are outbidding museums for T. rex fossils, raising concerns about the accessibility of scientific research.

AI/ML arXiv cs.AI

OOD-RL-Bench: A Benchmark Framework for Out-of-Distribution Detection in Reinforcement Learning

Introduction of OOD-RL-Bench, a framework for detecting out-of-distribution conditions in reinforcement learning trajectories.

AI/ML arXiv cs.AI

Mind the Gap: Promises and Pitfalls of Hierarchical Planning in LeWorldModel

Research on Hi-LeWM, an extension of LeWorldModel using hierarchical planning to improve long-horizon goal-conditioned control.

AI/ML arXiv cs.AI

Traceback Translators Against Forgetting in Continual Fake Speech Detection

A forgetting-resilient solution for continual fake speech detection using domain translators to remap new feature spaces.

AI/ML arXiv cs.AI

Deep Learning-based Surrogate Modelling of the LOD Method for Multiscale Problems

Introduction of LOD-MSNO, a hybrid neural operator that uses the Localized Orthogonal Decomposition method as a prior for multiscale PDEs.