All Articles
16276 articles total
HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization
HeraSys is introduced as an LLM serving system that optimizes concurrent agentic workflows through structural node merging and load-aware scheduling.
Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization
A new hardware-aware multi-objective structured pruning framework for LLMs to optimize for latency and size on edge devices.
Source-Aware Reranking for Retrieval-Augmented Generation: A Reliability Prior Approach
An evaluation of source-aware reranking in RAG pipelines that incorporates source reliability priors to improve retrieval precision.
The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation
Research highlighting the 'scaffold effect', where the choice of agent harness significantly impacts the token efficiency and performance of coding agents.
MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models
Introduction of MM-ShiftKV, a training-free KV selection method for Multimodal LLMs to reduce memory footprint during prefill.
Donate to GrapheneOS
A request for donations to support GrapheneOS, a privacy and security-focused mobile operating system.
Ozlo’s Sleepbuds 2 build on Bose’s sleep earbud legacy
Ozlo releases Sleepbuds 2, updating the sleep earbud product line originally started by Bose with better battery and connectivity.
The robot NASA hired to lift a orbital telescope is tumbling out of control
A NASA orbital telescope robot is malfunctioning due to reaction wheel failures and thruster problems, causing it to tumble.
Is it illegal to trick the US government into wiping your phone during a questionably legal search?
A Georgia man faces felony charges for wiping his phone during a CBP search, raising questions about the legality of data destruction during government questioning.
AI’s finally expensive enough to make Wall Street nervous
Investors are becoming concerned as Google increases its AI spending estimates, highlighting the difficulty of forecasting costs in the AI race.
This comfy gaming headset that can play audio from two sources is $25
The EPOS H3 Hybrid gaming headset is on a deep discount at Woot, offering dual audio sources and Bluetooth connectivity.
Logitech will pull a Nintendo — only European mice will come with replaceable batteries
Logitech plans to offer user-replaceable batteries in its wireless mice only for the European market to comply with upcoming EU laws.
Instacart's CTO says AI made the company stop worrying about tech debt
Instacart's CTO discusses how AI agents are now handling the majority of code generation and SRE tasks, reducing the burden of tech debt.
Too much evidence, too little time: From text to actionable recommendations through multi-objective evidence reasoning
Researchers introduce SCEPTER, a framework that uses LLMs and PubMedBERT to transform clinical case descriptions into evidence-based medical recommendations.
Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models
A study explores how long-context LLMs use a one-dimensional temporal code to mimic human episodic-like order memory.
Deflock Casa Grande
This entry refers to a discussion on Hacker News regarding 'Deflock Casa Grande', but provides no detailed content.
MCP 2026-07-28 Specification: transport going stateless
An update regarding the MCP 2026-07-28 Specification, focusing on moving the transport layer to a stateless model.
Waymo, robotaxi operators face fresh scrutiny over emergency response failures
Waymo and other robotaxi operators are facing federal scrutiny over safety standards after incidents where vehicles blocked emergency responders.
Apple won’t turn on any ‘restricted mode’ for missed lease payments
Apple clarifies that devices leased through its Upgrade program will not enter a 'restricted mode' if payments are missed, despite evidence of such code in iOS betas.
Keyword Matters: Unveiling the Energy Sensitivity of On-Device LLM Prompting
Research demonstrating that prompt engineering can be used as a lightweight method to reduce energy consumption for on-device LLM inference.