All Articles
17661 articles total
Kimi K2.7 Code is generally available in GitHub Copilot
Kimi K2.7 Code, a coding-specific model, is now available for use within GitHub Copilot.
Avo 4 released – 15 months and 2000 commits later
Avo 4 has been released after 15 months of development and 2000 commits.
A new Android malware from Google
Reports of a new Android malware originating from Google have emerged.
Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
Researchers introduce OpenAgent to study the fragility of static training in LLM tool-use agents and propose Perturbation-Augmented Fine-Tuning to improve robustness.
Optimal Resource Utilization for Autonomous Laboratory Orchestrators
A new two-step method combining constraint programming and status dependencies is proposed to optimize resource utilization in autonomous laboratory orchestrators.
Theoria: Rewrite-Acceptability Verification over Informal Reasoning States
Theoria is a new verification architecture that converts AI solutions into auditable typed state transitions to ensure rewrite-acceptability and reduce hidden premises.
AutoMem: Automated Learning of Memory as a Cognitive Skill
AutoMem is a framework that treats memory management as a trainable cognitive skill for LLMs, significantly improving performance on long-horizon tasks.
UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios
UltraFlux is a data-model co-design approach for native 4K text-to-image generation, utilizing a specialized corpus and architectural improvements for high-fidelity output.
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
DigitalCoach introduces a multimodal dataset of human expert-novice coaching sessions to evaluate and improve how AI agents teach humans to use software.
From "Strings" to "Things" for Personal Knowledge Graphs: Evaluating LLM Triple Extraction for Recommendation Systems
A study evaluates the use of lightweight LLMs (Qwen, Gemma) for extracting RDF-compliant triples from conversational data to build Personal Knowledge Graphs for recommendations.
Indian tech tycoon bets $30M of his own money to build AI alternative to Microsoft Office
Bhavin Turakhia is investing $30M to develop 'Neo', an AI-driven alternative to Microsoft Office and Google Apps.
AGI Maze as a Benchmark Framework for World-Modeling Agents
Researchers introduce AGI Maze, a benchmark framework designed to test if LLM agents can build and maintain persistent internal representations of a world state.
Coachable agents for interactive gameplay
A new framework for 'coachable agents' allows real-time control over the behavioral styles of AI agents in complex domains like video games and robotics.
Self-GC: Self-Governing Context for Long-Horizon LLM Agents
Self-GC introduces a context management system for long-horizon LLM agents that treats context as indexed, recoverable objects rather than simple text.
Self-Evolving Agents with Anytime-Valid Certificates
SEA (Self-Evolving Agents) uses a steering adapter and anytime-valid certificates to allow agents to self-modify without causing performance regressions.
Two AI Metrics Diverged: Will it Make All the Difference?
A study explores whether AI capabilities will concentrate among a few wealthy actors or proliferate to smaller models based on the nature of performance metrics.
Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination
Graph-PRefLexOR uses graph-native reinforcement learning to create traceable and interpretable scientific hypotheses in materials science.
Bayesian Uncertainty Propagation for Agentic RAG Pipelines: A Proof-of-Concept Study on Multi-Hop Question Answering
A proof-of-concept study proposes using Bayesian networks to propagate uncertainty signals in Agentic RAG pipelines to identify failure points.
PedNStream: Scalable Network Flow Simulation for Pedestrian Traffic Management
PedNStream is an open-source Python-native simulator for macroscopic pedestrian network flow and traffic management.
Agentic generation of verifiable rules for deterministic, self-expanding reaction classification
A multi-agent LLM pipeline automatically classifies chemical reactions and generates verifiable symbolic rules, expanding reaction taxonomies without human curation.