AI/ML arXiv cs.AI

Learning What to Remember: Observability-Safe Memory Retention via Constrained Optimization for Long-Horizon Language Agents

The OSL-MR framework treats memory retention in long-horizon language agents as a constrained stochastic optimization problem to improve resource allocation and observability.

AI/ML arXiv cs.AI

MoCA-Agent: A Market-of-Claims Code Agent for Financial and Numerical Reasoning

MoCA-Agent is a code agent for financial reasoning that decomposes questions into atomic claims and uses a market-like verification process to synthesize executable Python programs.

AI/ML arXiv cs.AI

Wisdom of Committee: Diverse Distillation from Large Foundation Models and Domain Experts

DiverseDistill is an interactive distillation framework that uses a committee of foundation models and domain experts to improve the performance of compact domain-specific models.

AI/ML arXiv cs.AI

Global Ease of Living Index: a machine learning framework for longitudinal analysis of major economies

The Global Ease of Living Index uses a machine learning framework and PCA/Factor Analysis to quantify quality of life across major economies since 1970.

AI/ML arXiv cs.AI

Simulation of Language Evolution under Regulated Social Media Platforms: A Synergistic Approach of Large Language Models and Genetic Algorithms

A multi-agent framework using LLMs and Genetic Algorithms is used to simulate how language evolves to evade moderation policies on social media platforms.

Tech Business/VC Hacker News

Americans express unease over SpaceX's influence on retirement savings

Americans express concern over SpaceX's growing influence on personal retirement savings.

Other Hacker News

How do flocking birds and schools of fish move?

An exploration into the biological mechanisms of how birds and fish coordinate movements in flocks and schools.

AI/ML Hacker News

A Perceptron in Age of Empires II

A technical demonstration of implementing a Perceptron neural network within the Age of Empires II game engine.

Cybersecurity Hacker News

AURpocalypse now: a look at the recent AUR attacks

An analysis of recent security attacks targeting the Arch User Repository (AUR).

AI/ML arXiv cs.AI

Mitigating Legibility Tax with Decoupled Prover-Verifier Games

Introduces Decoupled Prover-Verifier Games to reduce the 'legibility tax' when making LLM outputs checkable by smaller models.

AI/ML arXiv cs.AI

PrototypeNAS: Rapid Design of Deep Neural Networks for Microcontroller Units

Presents PrototypeNAS, a zero-shot neural architecture search method for rapidly designing DNNs for microcontrollers.

AI/ML arXiv cs.AI

The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation

Identifies the 'scaffold effect' where clinical VLMs show fake performance gains based on prompt framing rather than actual data integration.

AI/ML arXiv cs.AI

CareTransition-Audit: A Benchmark to Audit Discharge Summaries for Efficient Care Transitions

Proposes CareTransition-Audit, an LLM-based framework for auditing hospital discharge summaries to improve patient safety.

AI/ML arXiv cs.AI

Too long; didn't solve

Research indicating that both prompt and solution length correlate with increased failure rates in LLM mathematical reasoning.

AI/ML arXiv cs.AI

CogniFold: Always-On Proactive Memory via Cognitive Folding

Introduces CogniFold, a brain-inspired proactive agent memory system using graph-topology self-organization.

Other Hacker News

Think of the Children: How to Force Real ID for All Internet Traffic (2023)

A discussion on the technical and social implications of forcing Real ID requirements for all internet traffic.

Software Engineering Hacker News

Surprising Economics of Load-Balanced Systems

An exploration of the unexpected economic and performance characteristics of load-balanced systems.

Software Engineering Hacker News

RhinoCollab a plugin for real-time editing for Rhino 3D

RhinoCollab introduces real-time collaborative editing capabilities to Rhino 3D via a plugin.

Cybersecurity Hacker News

Hide Secrets from AI Agents and NPM install using Airgap

Airgap is presented as a tool to prevent secrets from being leaked to AI agents and during NPM installations.

Cybersecurity TechCrunch

Encryption, spyware, and now Mythos: History shows why cyber export control doesn’t work

An analysis of why cybersecurity export controls, including those for AI models like Anthropic's Mythos, are historically ineffective.