AI/ML arXiv cs.AI

PRESTO: Prefix-Aligned Tree Drafting for Diffusion Speculative Decoding

A new framework called PRESTO enables more efficient speculative decoding for diffusion models by using prefix-aligned tree drafting.

AI/ML arXiv cs.AI

CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants

A new benchmark called CallBench introduces a way to evaluate dual-goal coordination in phone call assistants.

AI/ML arXiv cs.AI

Answering Path Queries under Linear and Guarded Existential Rules

A study on the complexity of answering path queries in knowledge bases governed by guarded existential rules.

AI/ML arXiv cs.AI

Fast Cross-Scenario Adaptation of CSI Models via Channel Conditional Parameter Generation

Researchers introduce CCPG, a method for rapid cross-scenario adaptation of CSI models in wireless communications using lightweight LoRA weights.

AI/ML arXiv cs.AI

TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs

A new training curriculum called TRACE improves tool retrieval in enterprise LLMs by grounding reasoning in business rules.

AI/ML Hacker News

Running Kimi K3 on a M1 Mac

Discussion on the process and possibility of running the Kimi K3 model on M1 Mac hardware.

Cybersecurity Hacker News

Anthropic publishes a practical key-recovery attack on HAWK-256

Anthropic shares research on a practical key-recovery attack targeting the HAWK-256 cryptographic algorithm.

Open Source Hacker News

Robotics development made dead simple (open source)

An open-source project aimed at simplifying the development process for robotics.

Software Engineering Hacker News

Toolcraft

An article or project titled 'Toolcraft', likely focusing on developer tooling.

Cybersecurity VentureBeat

Visa used Mythos to hunt for bugs in its own payment network, then open-sourced the harness that made it possible

Visa has open-sourced its 'Vulnerability Agentic Harness', a tool used with AI models like Claude Mythos to automate the discovery and remediation of complex exploit chains.

AI/ML arXiv cs.AI

Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning

Introduction of Opti-Q, a cost-based optimizer that plans the execution of multi-LLM orchestrations to balance quality, cost, latency, and energy.

AI/ML arXiv cs.AI

CHS-SQL: A Text-to-SQL approach based on Confidence-Guided Heuristic Search Schema Linking process

CHS-SQL presents a novel Text-to-SQL framework for Small Language Models that optimizes the schema linking process using heuristic search and model confidence.

AI/ML arXiv cs.AI

TokenMem: Faithful Knowledge Injection for Frozen LLMs

TokenMem is a lightweight memory system that uses a cross-attention channel to inject knowledge into frozen LLMs, reducing conflicts between external and parametric memory.

AI/ML arXiv cs.AI

Masked Distillation: Internalizing the Chain-of-Thought in Language Models

The 'masked distillation' framework allows student LLMs to internalize the chain-of-thought reasoning of teacher models to produce answers more efficiently.

AI/ML arXiv cs.AI

VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing

VlogReward is a reward model designed to provide multi-dimensional evaluation and feedback for automated vlog editing using an enhanced GRPO framework.

Other Hacker News

Half-Life ported to Mac OS 9

The classic game Half-Life has been ported to Mac OS 9, bringing a legendary FPS experience to legacy Apple hardware.

Tech Business/VC TechCrunch

Bot-detection startup Spur nabs $200M from Insight

Bot-detection startup Spur Intelligence raised $200 million in funding from Insight Partners to improve human vs. bot traffic identification.

Cybersecurity The Verge

Ariana Grande is suing the hackers who’ve been leaking her songs and videos for years

Pop star Ariana Grande is suing unidentified hackers for stealing and leaking dozens of unreleased songs and videos.

AI/ML arXiv cs.AI

Chart Deception in Vision-Language Models: From Vulnerability to Mitigation

Researchers introduce VisDeception, a benchmark to evaluate how Vision-Language Models (VLMs) are misled by deceptive chart designs, and a multi-agent mitigation framework.

AI/ML arXiv cs.AI

DeepLook: Deeper Thinking with Lookahead

DeepLook is a training-free monitor-and-intervene framework that improves LLM reasoning by concentrating lookahead compute at uncertainty bottlenecks.