Portrait

Yufa Zhou

Logo CS PhD Student @ Duke

About Me

I am a CS PhD student at Duke University, advised by Prof. Anru Zhang, interning at ByteDance, doing autoresearch on post-training.

I also run Nonlinearity, an independent research effort on self-influencing systems.

Education

  • Duke University
    Duke University
    Ph.D. in Computer Science
    Aug. 2025 – Present
  • University of Pennsylvania
    University of Pennsylvania
    M.S.E. in Scientific Computing
    Aug. 2023 - May. 2025
  • Wuhan University
    Wuhan University
    B.E. in Engineering Mechanics
    Sep. 2019 - Jul. 2023

Experience

  • ByteDance, San Jose, CA
    ByteDance, San Jose, CA
    Research Scientist Intern
    Jun. 2026 - Present

News

2026
5 papers accepted at ICLR (×3), ICML, and NeurIPS 2026
Sep 24
Accepted the Research Scientist Intern offer at ByteDance, San Jose, CA
Jan 16
2025
6 papers accepted at AAAI (×2), ICLR, AISTATS, ICCV, and NeurIPS 2025, plus a NeurIPS workshop oral
Sep 23
Accepted the Ph.D. offer in Computer Science at Duke University
Feb 27

Selected Publications (view all )

The Geometry of Contextual Relations: Language Models Address Facts by Order of Mention

The Geometry of Contextual Relations: Language Models Address Facts by Order of Mention

Yufa Zhou

arXiv preprint 2026

We show that LLMs reach an in-context fact by its order of mention, not by the name in the question: a shared, steerable, low-rank subspace of fact addresses, found across 14 models from 1.5B to 32B parameters and formed early in pretraining, lets a question point to the fact it asks about while the context supplies what that fact says.

×
BibTeX Citation
@article{zhou2026geometry, title = {The Geometry of Contextual Relations: Language Models Address Facts by Order of Mention}, author = {Zhou, Yufa}, journal = {arXiv preprint arXiv:2610.00910}, year = {2026} }

The Geometry of Contextual Relations: Language Models Address Facts by Order of Mention

Yufa Zhou

arXiv preprint 2026

We show that LLMs reach an in-context fact by its order of mention, not by the name in the question: a shared, steerable, low-rank subspace of fact addresses, found across 14 models from 1.5B to 32B parameters and formed early in pretraining, lets a question point to the fact it asks about while the context supplies what that fact says.

×
BibTeX Citation
@article{zhou2026geometry, title = {The Geometry of Contextual Relations: Language Models Address Facts by Order of Mention}, author = {Zhou, Yufa}, journal = {arXiv preprint arXiv:2610.00910}, year = {2026} }
The Geometry of Reasoning: Flowing Logics in Representation Space

The Geometry of Reasoning: Flowing Logics in Representation Space

Yufa Zhou*, Yixiao Wang*, Xunjian Yin*, Shuyan Zhou, Anru R. Zhang(* equal contribution)

ICLR 2026

We study how LLMs “think” through their embeddings by introducing a geometric framework of reasoning flows, where reasoning emerges as smooth trajectories in representation space whose velocity and curvature are governed by logical structure rather than surface semantics, validated through cross-topic and cross-language experiments, opening a new lens for interpretability.

×
BibTeX Citation
@inproceedings{zhou2025geometry, title = {The Geometry of Reasoning: Flowing Logics in Representation Space}, author = {Zhou, Yufa and Wang, Yixiao and Yin, Xunjian and Zhou, Shuyan and Zhang, Anru R.}, booktitle = {The Fourteenth International Conference on Learning Representations}, year = {2026}, url = {https://openreview.net/forum?id=ixr5Pcabq7} }

The Geometry of Reasoning: Flowing Logics in Representation Space

Yufa Zhou*, Yixiao Wang*, Xunjian Yin*, Shuyan Zhou, Anru R. Zhang(* equal contribution)

ICLR 2026

We study how LLMs “think” through their embeddings by introducing a geometric framework of reasoning flows, where reasoning emerges as smooth trajectories in representation space whose velocity and curvature are governed by logical structure rather than surface semantics, validated through cross-topic and cross-language experiments, opening a new lens for interpretability.

×
BibTeX Citation
@inproceedings{zhou2025geometry, title = {The Geometry of Reasoning: Flowing Logics in Representation Space}, author = {Zhou, Yufa and Wang, Yixiao and Yin, Xunjian and Zhou, Shuyan and Zhang, Anru R.}, booktitle = {The Fourteenth International Conference on Learning Representations}, year = {2026}, url = {https://openreview.net/forum?id=ixr5Pcabq7} }
Why Do Transformers Fail to Forecast Time Series In-Context?

Why Do Transformers Fail to Forecast Time Series In-Context?

Yufa Zhou*, Yixiao Wang*, Surbhi Goel, Anru R. Zhang(* equal contribution)

NeurIPS 2025 Workshop: What Can('t) Transformers Do? Oral (3/68 ≈ 4.4%)

We analyze why Transformers fail in time-series forecasting through in-context learning theory, proving that, under AR($p$) data, linear self-attention cannot outperform classical linear predictors and suffers a strict $O(1/n)$ excess-risk gap, while chain-of-thought inference compounds errors exponentially—revealing fundamental representational limits of attention and offering principled insights.

×
BibTeX Citation
@article{zhou2025tsf, title={Why Do Transformers Fail to Forecast Time Series In-Context?}, author={Zhou, Yufa and Wang, Yixiao and Goel, Surbhi and Zhang, Anru R.}, journal={arXiv preprint arXiv:2510.09776}, year={2025} }

Why Do Transformers Fail to Forecast Time Series In-Context?

Yufa Zhou*, Yixiao Wang*, Surbhi Goel, Anru R. Zhang(* equal contribution)

NeurIPS 2025 Workshop: What Can('t) Transformers Do? Oral (3/68 ≈ 4.4%)

We analyze why Transformers fail in time-series forecasting through in-context learning theory, proving that, under AR($p$) data, linear self-attention cannot outperform classical linear predictors and suffers a strict $O(1/n)$ excess-risk gap, while chain-of-thought inference compounds errors exponentially—revealing fundamental representational limits of attention and offering principled insights.

×
BibTeX Citation
@article{zhou2025tsf, title={Why Do Transformers Fail to Forecast Time Series In-Context?}, author={Zhou, Yufa and Wang, Yixiao and Goel, Surbhi and Zhang, Anru R.}, journal={arXiv preprint arXiv:2510.09776}, year={2025} }

Mentees

Teaching

Academic Services

  • Conference Reviewer: ICLR (2025, 2026), NeurIPS 2026, ICML 2026, AAAI (2026, 2027).