All Articles
17646 articles total
A Practice Auditing Framework for Large Language Model Use: Collective Empiricism, Pseudo-Rational Cognition, and Governance of AI-Generated Content
The paper proposes a practice auditing framework to govern AI-generated content and mitigate risks like 'pseudo-rational cognition' and memory pollution.
Structuring the Space of Sociotechnical Alignment
This research argues for a more systematic, human-centered framework to specify and evaluate 'social desirability' in sociotechnical AI alignment.
Collaborative Disagreement Resolution for Scalable Oversight
Researchers propose 'disagreement resolution,' a collaborative truth-seeking paradigm for AI oversight that outperforms adversarial debate in judging accuracy.
How Indian Dermatologists are Utilizing Artificial Intelligence for Clinical Practice and Workflow Management: A Nationwide Survey with a Special Focus on atopic dermatitis
A survey of Indian dermatologists reveals that general-purpose AI is mainly used for administrative tasks rather than specialized clinical diagnostic workflows.
Beyond Detection: Redesigning Assessment and Governande of Generative AI at the Universidad Polit\'ecnica de Madrid (UPM)
The Universidad Politécnica de Madrid proposes a strategic framework to move beyond AI detection toward deep pedagogical adoption and AI literacy.
Wordgard: The new in-browser rich-text editor from the creator of ProseMirror
The creator of ProseMirror has released Wordgard, a new rich-text editor designed for use in the browser.
Q&A with Micron's VP and GM of Memory
A Q&A session featuring Micron's VP and GM of Memory discussing memory technology trends.
ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
RECONTEXT is a training-free inference method that improves long-context reasoning in LLMs by recursively replaying relevant evidence.
Online Safety Monitoring for LLMs
Researchers propose a real-time safety monitor for LLMs that uses a verifier signal from an external model to trigger alarms based on risk-controlled thresholds.
Distributed Attacks in Persistent-State AI Control
Study on 'Iterative VibeCoding' reveals how AI agents can distribute malicious payloads across multiple pull requests in persistent codebases to evade detection.
TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language Models
TokenScope is an interactive interpretability tool for decoder-based LLMs that provides token-level metrics and structural analysis during code generation.
Safeguarding LLM Agents from Misalignment through Provenance Analysis
ProvenanceGuard is a multi-stage pipeline that prevents LLM agent misalignment by verifying if tool calls are supported by traceable evidence in the context.
Kara: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression
Kara is a sliding-window KV cache compression method that reduces memory usage and increases throughput for reasoning LLMs, integrated into the KvLLM framework.
SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings
SPARCLE is a speaker-aware grapheme representation model that improves text-to-speech quality in low-resource settings by aligning graphemes with acoustic representations.
Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment
Research shows that BPE tokenization creates gaps in LLM alignment, allowing attackers to bypass safety filters by fragmenting safety-critical words.
Half-Baked Product
A Hacker News discussion regarding the concept of a 'Half-Baked Product'.
DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models
Introduces DRIFTLENS, a framework to measure 'reasoning drift' in personalized LLMs where user-specific memory alters the model's reasoning trajectory.
Hardware-Enforced Semantic Coordination for Safety-Critical Real-Time Autonomous Systems
Proposes a hardware-enforced semantic coordination architecture using FPGAs to ensure deterministic safety and timing in autonomous AI systems.
Steerability via constraints: a substrate for scalable oversight of coding agents
Argues for using traditional engineering constraints (access control, coding conventions) to improve the oversight and safety of coding agents.
Fast Multi-dimensional Refusal Subspaces via RFM-AGOP
Presents RFM-AGOP, a method for rapidly identifying multi-dimensional refusal subspaces in LLMs to improve safety and interpretability.