NOW · UPDATED AUGUST 26, 2026
What I am paying attention to now.
This is a living page: less formal than a CV, more current than a biography, and intentionally focused on questions rather than announcements.
BUILDING
Memory-aware systems for Audio LLMs
Exploring how semantic and acoustic evidence can guide KV-cache retention instead of relying on uniform compression.
THINKING
When diffusion parallelism becomes real speed
Looking beyond theoretical parallel decoding toward strategies whose gains survive scheduling, verification, and wall-clock measurement.
WRITING
A quieter record beside the research page
Using the Journal for travel, photographs, reading, and thoughts that do not need to become a polished research claim.
OPEN QUESTIONS
Questions worth keeping visible
-
01
Can cache compression become semantic-aware without introducing a second expensive model?
-
02
Which diffusion-LLM acceleration gains remain after end-to-end latency and hardware utilization are counted?
-
03
How should long-context audio benchmarks balance understanding quality, memory footprint, and response latency?
Maintenance note: update this page directly in _pages/now.md whenever the center of gravity changes.