Other Hacker News

John Carmack on Fabrice Bellard

A discussion on Hacker News featuring perspectives from John Carmack regarding the work of Fabrice Bellard.

Other Hacker News

Chili peppers of the world: cultivars, species, and heat

A community discussion about the various cultivars and heat levels of chili peppers worldwide.

AI/ML arXiv cs.AI

QoS-Aware Token Scheduling and Private Data Valuation for Multi-Modal Agentic Networks

Research on fair token allocation and private data valuation in decentralized multi-modal agentic networks using differentially private prototypes.

AI/ML arXiv cs.AI

Do we have the knowledge we need? Rethinking human-AI decision-making in corporations

A position paper exploring how organizations should maintain knowledge for both human and AI accessibility and how to allocate agency between them.

AI/ML arXiv cs.AI

Large Language Models as Optimizers: A Survey of Direct vs. Tool-Augmented Approaches and Their Performance Frontiers

A survey of LLMs as optimizers, analyzing direct, tool-augmented, and tool-creating optimization paradigms.

AI/ML arXiv cs.AI

Your Agent Has a Genome: Sequence-Level Behavioral Analysis and Runtime Governance of LLM-Powered Autonomous Agents

Introduces Base Sequence Analysis for LLM agent behavior and Governor, a runtime intervention system that increases success rates and reduces token costs.

AI/ML arXiv cs.AI

Agentic Retrieval and Reinforcement Learned Equation Chains: A Controlled Generation Framework for Complex and Novel Physics Word Problems

Introduces ARVRE, a framework using reinforcement learning and agentic RAG to generate mathematically valid and complex physics word problems.

AI/ML arXiv cs.AI

Integrating Reasoning and Generalization in Text-to-SQL via Self-Enhanced Fine-Tuning

Presents CoTE-SQL, a method for enhancing Text-to-SQL generation using self-enhanced reasoning traces and error-aware revision.

AI/ML arXiv cs.AI

NeuroSymbolic AI for Legal AI-TRISM: Trustworthy, Reliable, Interpretable, Safe Models

Proposed TRISM framework combining NeuroSymbolic AI with LLMs to improve trust, reliability, and interpretability in legal AI applications.

AI/ML arXiv cs.AI

Towards Next-Generation Healthcare: A Survey of Medical Embodied AI for Perception, Decision-Making, and Action

A comprehensive survey on Medical Embodied AI, focusing on the integration of perception, decision-making, and action in clinical environments.

Cybersecurity Hacker News

I Could've Rickrolled the FIFA World Cup. All I Needed Was My ID

A first-person account of a potential security vulnerability at the FIFA World Cup involving ID badge access.

AI/ML arXiv cs.AI

Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds

Research on reward hacking in LLM agents using a text-based evaluation suite to show that proxy-reward failures resist standard mitigations.

AI/ML arXiv cs.AI

Hierarchical Modeling of ICD Codes in EHR Foundation Models

A study on improving EHR foundation models by explicitly incorporating the hierarchical structure of ICD-10-CM diagnosis codes.

AI/ML arXiv cs.AI

Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines

Proposed method for anytime-valid attribution in LLM evaluation pipelines to distinguish between product drift and judge model drift.

AI/ML arXiv cs.AI

Towards End-to-End Automation of AI Research

Introduction of 'The AI Scientist', an end-to-end automated system that can generate research ideas, execute experiments, and write scientific manuscripts.

AI/ML arXiv cs.AI

Synthetic Counteradaptation: A Principle of Human-AI Co-evolution

Introduction of the 'synthetic counteradaptation' principle to describe the recursive co-evolution of strategies between humans and AI.

AI/ML arXiv cs.AI

Toward Vibe Medicine: A Self-Evolving Multi-Agent Framework for Clinical Decision Support

VIBEMed is presented as a self-evolving multi-agent framework for clinical decision support that learns from patient outcomes and past failures.

AI/ML arXiv cs.AI

Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning

A mechanistic interpretability audit of LLaMA 3.1-8B-Instruct's ethical reasoning, revealing that surface-level prompts often dominate the internal computation.

AI/ML arXiv cs.AI

ToolMenuBench: Benchmarking Tool-Menu Filtering Strategies for Reliable and Efficient LLM Agents

ToolMenuBench is a new benchmark for evaluating how the selection and filtering of tools provided to LLM agents affects reliability and efficiency.

AI/ML arXiv cs.AI

Minimal Oversight: Uncertainty-Aware Governance for Delegated AI Systems

A framework for uncertainty-aware governance in delegated AI systems using the Minimum Sufficient Oversight Principle (MSO).