About

Hi — I’m Zeman Li (李泽慢), a third-year Ph.D. student in Industrial & Systems Engineering at the University of Southern California, advised by Prof. Meisam Razaviyayn. I am currently a Student Researcher at Google Research.

Before USC, I earned a B.S. in Computer Engineering at Georgia Tech, where I was fortunate to work with Prof. Yao Xie, and a B.S. in Applied Mathematics at Emory University.

Research interests

I work at the intersection of large-scale optimization and machine learning: data-efficient methods for training foundation models, test-time / context memorization, and privacy-preserving learning. Recent work has appeared at ICLR, ICML, and NeurIPS.

Contact

News

  • Paper accepted at ICLR 2026 (Spotlight): TNT — Improving Chunkwise Training for Test-Time Memorization.
  • Paper accepted at NeurIPS 2026 (Spotlight): PiKE — Adaptive Data Mixing for Multi-Task Learning Under Low Gradient Conflicts.
  • Paper accepted at ICML 2025: Synthetic Text Generation for Training Large Language Models via Gradient Matching.
  • Paper accepted at ICLR 2025: Addax — Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models.
  • Paper accepted at ICML 2024: Optimal Differentially Private Learning with Public Data.
  • Started Ph.D. at USC ISE, advised by Prof. Meisam Razaviyayn.

Publications

Auto-synced from Semantic Scholar — last build . ⊕ Atom

9 entries, newest first.

  1. Mem3R: Streaming 3D Reconstruction with Hybrid Memory via Test-Time Training
    arXiv (preprint) · · paper
  2. Memory Caching: RNNs with Growing Memory
    arXiv (preprint) · · paper
  3. Online Neural Space Time Memory for Dynamic Novel View Synthesis
    arXiv (preprint) · · paper
  4. PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts
    NeurIPS 2026 (Spotlight) · · paper
  5. TNT: Improving Chunkwise Training for Test-Time Memorization
    ICLR 2026 (Spotlight) · · paper
  6. Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
    ICLR 2025 · · paper
  7. ATLAS: Learning to Optimally Memorize the Context at Test Time
    arXiv (preprint) · · paper
  8. Synthetic Text Generation for Training Large Language Models via Gradient Matching
    ICML 2025 · · paper
  9. Optimal Differentially Private Model Training with Public Data
    ICML 2024 · · paper

Vita

Education

  • Ph.D., Industrial & Systems Engineering, University of Southern California
    Advisor: Prof. Meisam Razaviyayn
  • B.S., Computer Engineering, Georgia Institute of Technology
    Advisor: Prof. Yao Xie
  • B.S., Applied Mathematics, Emory University

Awards

Knowledge Base

Working notes I keep while reading — published openly in case someone else finds them useful.

Playground

Some small interactive things I’ve built into this site for fun.