AI/ML Hacker News

Kimi K2.7 Code is generally available in GitHub Copilot

Kimi K2.7 Code, a coding-specific model, is now available for use within GitHub Copilot.

Software Engineering Hacker News

Avo 4 released – 15 months and 2000 commits later

Avo 4 has been released after 15 months of development and 2000 commits.

Cybersecurity Hacker News

A new Android malware from Google

Reports of a new Android malware originating from Google have emerged.

AI/ML arXiv cs.AI

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

Researchers introduce OpenAgent to study the fragility of static training in LLM tool-use agents and propose Perturbation-Augmented Fine-Tuning to improve robustness.

AI/ML arXiv cs.AI

Optimal Resource Utilization for Autonomous Laboratory Orchestrators

A new two-step method combining constraint programming and status dependencies is proposed to optimize resource utilization in autonomous laboratory orchestrators.

AI/ML arXiv cs.AI

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States

Theoria is a new verification architecture that converts AI solutions into auditable typed state transitions to ensure rewrite-acceptability and reduce hidden premises.

AI/ML arXiv cs.AI

AutoMem: Automated Learning of Memory as a Cognitive Skill

AutoMem is a framework that treats memory management as a trainable cognitive skill for LLMs, significantly improving performance on long-horizon tasks.

AI/ML arXiv cs.AI

UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios

UltraFlux is a data-model co-design approach for native 4K text-to-image generation, utilizing a specialized corpus and architectural improvements for high-fidelity output.

AI/ML arXiv cs.AI

DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching

DigitalCoach introduces a multimodal dataset of human expert-novice coaching sessions to evaluate and improve how AI agents teach humans to use software.

AI/ML arXiv cs.AI

From "Strings" to "Things" for Personal Knowledge Graphs: Evaluating LLM Triple Extraction for Recommendation Systems

A study evaluates the use of lightweight LLMs (Qwen, Gemma) for extracting RDF-compliant triples from conversational data to build Personal Knowledge Graphs for recommendations.

Tech Business/VC TechCrunch

Indian tech tycoon bets $30M of his own money to build AI alternative to Microsoft Office

Bhavin Turakhia is investing $30M to develop 'Neo', an AI-driven alternative to Microsoft Office and Google Apps.

AI/ML arXiv cs.AI

AGI Maze as a Benchmark Framework for World-Modeling Agents

Researchers introduce AGI Maze, a benchmark framework designed to test if LLM agents can build and maintain persistent internal representations of a world state.

AI/ML arXiv cs.AI

Coachable agents for interactive gameplay

A new framework for 'coachable agents' allows real-time control over the behavioral styles of AI agents in complex domains like video games and robotics.

AI/ML arXiv cs.AI

Self-GC: Self-Governing Context for Long-Horizon LLM Agents

Self-GC introduces a context management system for long-horizon LLM agents that treats context as indexed, recoverable objects rather than simple text.

AI/ML arXiv cs.AI

Self-Evolving Agents with Anytime-Valid Certificates

SEA (Self-Evolving Agents) uses a steering adapter and anytime-valid certificates to allow agents to self-modify without causing performance regressions.

AI/ML arXiv cs.AI

Two AI Metrics Diverged: Will it Make All the Difference?

A study explores whether AI capabilities will concentrate among a few wealthy actors or proliferate to smaller models based on the nature of performance metrics.

AI/ML arXiv cs.AI

Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

Graph-PRefLexOR uses graph-native reinforcement learning to create traceable and interpretable scientific hypotheses in materials science.

AI/ML arXiv cs.AI

Bayesian Uncertainty Propagation for Agentic RAG Pipelines: A Proof-of-Concept Study on Multi-Hop Question Answering

A proof-of-concept study proposes using Bayesian networks to propagate uncertainty signals in Agentic RAG pipelines to identify failure points.

Open Source arXiv cs.AI

PedNStream: Scalable Network Flow Simulation for Pedestrian Traffic Management

PedNStream is an open-source Python-native simulator for macroscopic pedestrian network flow and traffic management.

AI/ML arXiv cs.AI

Agentic generation of verifiable rules for deterministic, self-expanding reaction classification

A multi-agent LLM pipeline automatically classifies chemical reactions and generates verifiable symbolic rules, expanding reaction taxonomies without human curation.