All Articles
16266 articles total
PRESTO: Prefix-Aligned Tree Drafting for Diffusion Speculative Decoding
A new framework called PRESTO enables more efficient speculative decoding for diffusion models by using prefix-aligned tree drafting.
CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants
A new benchmark called CallBench introduces a way to evaluate dual-goal coordination in phone call assistants.
Answering Path Queries under Linear and Guarded Existential Rules
A study on the complexity of answering path queries in knowledge bases governed by guarded existential rules.
Fast Cross-Scenario Adaptation of CSI Models via Channel Conditional Parameter Generation
Researchers introduce CCPG, a method for rapid cross-scenario adaptation of CSI models in wireless communications using lightweight LoRA weights.
TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs
A new training curriculum called TRACE improves tool retrieval in enterprise LLMs by grounding reasoning in business rules.
Running Kimi K3 on a M1 Mac
Discussion on the process and possibility of running the Kimi K3 model on M1 Mac hardware.
Anthropic publishes a practical key-recovery attack on HAWK-256
Anthropic shares research on a practical key-recovery attack targeting the HAWK-256 cryptographic algorithm.
Robotics development made dead simple (open source)
An open-source project aimed at simplifying the development process for robotics.
Toolcraft
An article or project titled 'Toolcraft', likely focusing on developer tooling.
Visa used Mythos to hunt for bugs in its own payment network, then open-sourced the harness that made it possible
Visa has open-sourced its 'Vulnerability Agentic Harness', a tool used with AI models like Claude Mythos to automate the discovery and remediation of complex exploit chains.
Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning
Introduction of Opti-Q, a cost-based optimizer that plans the execution of multi-LLM orchestrations to balance quality, cost, latency, and energy.
CHS-SQL: A Text-to-SQL approach based on Confidence-Guided Heuristic Search Schema Linking process
CHS-SQL presents a novel Text-to-SQL framework for Small Language Models that optimizes the schema linking process using heuristic search and model confidence.
TokenMem: Faithful Knowledge Injection for Frozen LLMs
TokenMem is a lightweight memory system that uses a cross-attention channel to inject knowledge into frozen LLMs, reducing conflicts between external and parametric memory.
Masked Distillation: Internalizing the Chain-of-Thought in Language Models
The 'masked distillation' framework allows student LLMs to internalize the chain-of-thought reasoning of teacher models to produce answers more efficiently.
VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing
VlogReward is a reward model designed to provide multi-dimensional evaluation and feedback for automated vlog editing using an enhanced GRPO framework.
Half-Life ported to Mac OS 9
The classic game Half-Life has been ported to Mac OS 9, bringing a legendary FPS experience to legacy Apple hardware.
Bot-detection startup Spur nabs $200M from Insight
Bot-detection startup Spur Intelligence raised $200 million in funding from Insight Partners to improve human vs. bot traffic identification.
Ariana Grande is suing the hackers who’ve been leaking her songs and videos for years
Pop star Ariana Grande is suing unidentified hackers for stealing and leaking dozens of unreleased songs and videos.
Chart Deception in Vision-Language Models: From Vulnerability to Mitigation
Researchers introduce VisDeception, a benchmark to evaluate how Vision-Language Models (VLMs) are misled by deceptive chart designs, and a multi-agent mitigation framework.
DeepLook: Deeper Thinking with Lookahead
DeepLook is a training-free monitor-and-intervene framework that improves LLM reasoning by concentrating lookahead compute at uncertainty bottlenecks.