PUMA Lab
  • Group
  • Publications
  • Research Blogs
  • Join
  • Lab Wiki

Research Blogs


OpenSafeIntent: A Diagnostic Benchmark for Intent-Calibrated Safe Completion

OpenSafeIntent: A Diagnostic Benchmark for Intent-Calibrated Safe Completion

Rheeya Uppaal

Today’s models lack safety consistency: the same dual-use task can flip from safe to unsafe with a change in intent or even wording.

R-KV: Redundancy-aware KV Cache Compression for Reasoning Models

R-KV: Redundancy-aware KV Cache Compression for Reasoning Models

Zefan Cai

Reasoning traces are far more redundant than they look: keeping just 10% of the KV cache can preserve essentially full reasoning accuracy and even improve it by filtering noise.

MENTOR: Efficient Multimodal‑Conditioned Tuning for Autoregressive Vision Generation Models

MENTOR: Efficient Multimodal‑Conditioned Tuning for Autoregressive Vision Generation Models

Zefan Cai

Autoregressive vision models can rival, and even surpass, diffusion models at multimodal image generation, while training on 10× less data and a fraction of the compute.

Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity

Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity

Rheeya Uppaal

DPO may be doing something much simpler than gradient descent suggests: learning a low-dimensional “toxic subspace” that can be removed directly with a single closed-form projection.

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Zefan Cai

PyramidKV trims the KV cache where attention matters least: cutting memory by ~88% without cutting model quality.

© 2026 PUMA Lab
Built with GitHub Pages + Jekyll