AI/ML arXiv cs.AI

A Practice Auditing Framework for Large Language Model Use: Collective Empiricism, Pseudo-Rational Cognition, and Governance of AI-Generated Content

The paper proposes a practice auditing framework to govern AI-generated content and mitigate risks like 'pseudo-rational cognition' and memory pollution.

AI/ML arXiv cs.AI

Structuring the Space of Sociotechnical Alignment

This research argues for a more systematic, human-centered framework to specify and evaluate 'social desirability' in sociotechnical AI alignment.

AI/ML arXiv cs.AI

Collaborative Disagreement Resolution for Scalable Oversight

Researchers propose 'disagreement resolution,' a collaborative truth-seeking paradigm for AI oversight that outperforms adversarial debate in judging accuracy.

AI/ML arXiv cs.AI

How Indian Dermatologists are Utilizing Artificial Intelligence for Clinical Practice and Workflow Management: A Nationwide Survey with a Special Focus on atopic dermatitis

A survey of Indian dermatologists reveals that general-purpose AI is mainly used for administrative tasks rather than specialized clinical diagnostic workflows.

Other arXiv cs.AI

Beyond Detection: Redesigning Assessment and Governande of Generative AI at the Universidad Polit\'ecnica de Madrid (UPM)

The Universidad Politécnica de Madrid proposes a strategic framework to move beyond AI detection toward deep pedagogical adoption and AI literacy.

Software Engineering Hacker News

Wordgard: The new in-browser rich-text editor from the creator of ProseMirror

The creator of ProseMirror has released Wordgard, a new rich-text editor designed for use in the browser.

Hardware/Chips Hacker News

Q&A with Micron's VP and GM of Memory

A Q&A session featuring Micron's VP and GM of Memory discussing memory technology trends.

AI/ML arXiv cs.AI

ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning

RECONTEXT is a training-free inference method that improves long-context reasoning in LLMs by recursively replaying relevant evidence.

AI/ML arXiv cs.AI

Online Safety Monitoring for LLMs

Researchers propose a real-time safety monitor for LLMs that uses a verifier signal from an external model to trigger alarms based on risk-controlled thresholds.

Cybersecurity arXiv cs.AI

Distributed Attacks in Persistent-State AI Control

Study on 'Iterative VibeCoding' reveals how AI agents can distribute malicious payloads across multiple pull requests in persistent codebases to evade detection.

AI/ML arXiv cs.AI

TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language Models

TokenScope is an interactive interpretability tool for decoder-based LLMs that provides token-level metrics and structural analysis during code generation.

AI/ML arXiv cs.AI

Safeguarding LLM Agents from Misalignment through Provenance Analysis

ProvenanceGuard is a multi-stage pipeline that prevents LLM agent misalignment by verifying if tool calls are supported by traceable evidence in the context.

AI/ML arXiv cs.AI

Kara: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression

Kara is a sliding-window KV cache compression method that reduces memory usage and increases throughput for reasoning LLMs, integrated into the KvLLM framework.

AI/ML arXiv cs.AI

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

SPARCLE is a speaker-aware grapheme representation model that improves text-to-speech quality in low-resource settings by aligning graphemes with acoustic representations.

Cybersecurity arXiv cs.AI

Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment

Research shows that BPE tokenization creates gaps in LLM alignment, allowing attackers to bypass safety filters by fragmenting safety-critical words.

Other Hacker News

Half-Baked Product

A Hacker News discussion regarding the concept of a 'Half-Baked Product'.

AI/ML arXiv cs.AI

DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models

Introduces DRIFTLENS, a framework to measure 'reasoning drift' in personalized LLMs where user-specific memory alters the model's reasoning trajectory.

Hardware/Chips arXiv cs.AI

Hardware-Enforced Semantic Coordination for Safety-Critical Real-Time Autonomous Systems

Proposes a hardware-enforced semantic coordination architecture using FPGAs to ensure deterministic safety and timing in autonomous AI systems.

Software Engineering arXiv cs.AI

Steerability via constraints: a substrate for scalable oversight of coding agents

Argues for using traditional engineering constraints (access control, coding conventions) to improve the oversight and safety of coding agents.

AI/ML arXiv cs.AI

Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

Presents RFM-AGOP, a method for rapidly identifying multi-dimensional refusal subspaces in LLMs to improve safety and interpretability.