NOW · UPDATED AUGUST 26, 2026

What I am paying attention to now.

This is a living page: less formal than a CV, more current than a biography, and intentionally focused on questions rather than announcements.

BUILDING

Memory-aware systems for Audio LLMs

Exploring how semantic and acoustic evidence can guide KV-cache retention instead of relying on uniform compression.

AudioKVKV CacheEfficient AI

THINKING

When diffusion parallelism becomes real speed

Looking beyond theoretical parallel decoding toward strategies whose gains survive scheduling, verification, and wall-clock measurement.

Diffusion LLMInferenceSystems

WRITING

A quieter record beside the research page

Using the Journal for travel, photographs, reading, and thoughts that do not need to become a polished research claim.

JournalNotesLife

OPEN QUESTIONS

Questions worth keeping visible

  1. 01

    Can cache compression become semantic-aware without introducing a second expensive model?

  2. 02

    Which diffusion-LLM acceleration gains remain after end-to-end latency and hardware utilization are counted?

  3. 03

    How should long-context audio benchmarks balance understanding quality, memory footprint, and response latency?

Maintenance note: update this page directly in _pages/now.md whenever the center of gravity changes.