BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//UM//UM*Events//EN
CALSCALE:GREGORIAN
BEGIN:VTIMEZONE
TZID:America/Detroit
TZURL:http://tzurl.org/zoneinfo/America/Detroit
X-LIC-LOCATION:America/Detroit
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20070311T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20071104T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20250203T095616
DTSTART;TZID=America/Detroit:20250211T160000
DTEND;TZID=America/Detroit:20250211T170000
SUMMARY:Workshop / Seminar:Statistics Department Seminar Series: Ye Tian\, PhD Candidate\, Department of Statistics\, Columbia University
DESCRIPTION:Abstract: Knowledge transfer\, a core human ability\, has inspired numerous data integration methods in machine learning and statistics. However\, data integration faces significant challenges: (1) unknown similarity between data sources\; (2) data contamination\; (3) high-dimensionality\; and (4) privacy constraints. This talk addresses these challenges in three parts across different contexts\, presenting both innovative statistical methodologies and theoretical insights.\n\nIn Part I\, I will introduce a transfer learning framework for high-dimensional generalized linear models that combines a pre-trained Lasso with a fine-tuning step. We provide theoretical guarantees for both estimation and inference.\n\nIn Part II\, I will explore an unsupervised learning setting where task-specific data is generated from a mixture model with heterogeneous mixture proportions. This complements the supervised learning setting discussed in Part I\, addressing scenarios where labeled data is unavailable. We propose a federated gradient EM algorithm that is communication-efficient and privacy-preserving\, providing estimation error bounds for the mixture model parameters.\n\nIn Part III\, I will introduce a representation-based multi-task learning framework that generalizes the distance-based similarity notion discussed in Parts I and II. This framework is closely related to modern applications of fine-tuning in image classification and natural language processing. I will discuss how this study enhances our understanding of the effectiveness of fine-tuning and the influence of data contamination on representation multi-task learning.\n\nFinally\, I will summarize the talk and briefly introduce my broader research interests. I will also discuss my work on imbalanced classification error control\, highlighting its role in improving classification performance for minority groups and its applications in biomedical and business contexts. The three main sections of this talk are based on a series of papers [TF23\, TWXF22\, TWF24\, TGF23] and a short course I co-taught at NESS 2024 [STL24]. The segment on imbalanced classification error control is based on our recent paper [TF24]. More about me and my research can be found at https://yet123.com.\n\n[TF23] Tian\, Y.\, & Feng\, Y. (2023). Transfer Learning under High-dimensional Generalized Linear Models. Journal of the American Statistical Association\, 118(544)\, 2684-2697.\n[TWXF22] Tian\, Y.\, Weng\, H.\, Xia\, L.\, & Feng\, Y. (2022). Unsupervised Multi-task and Transfer Learning on Gaussian Mixture Models. arXiv preprint arXiv:2209.15224. \n[TWF24] Tian\, Y.\, Weng\, H.\, & Feng\, Y. (2024). Towards the Theory of Unsupervised Federated Learning: Non-asymptotic Analysis of Federated EM Algorithms. ICML 2024. \n[TGF23] Tian\, Y.\, Gu\, Y.\, & Feng\, Y. (2023). Learning from Similar Linear Representations: Adaptivity\, Minimaxity\, and Robustness. arXiv preprint arXiv:2303.17765. \n[STL24] A (Selective) Introduction to the Statistics Foundations of Transfer Learning. (2024).\n[TF24] Tian\, Y.\, & Feng\, Y. (2024). Neyman-Pearson Multi-class Classification via Cost-sensitive Learning. Journal of the American Statistical Association\, 1-15. \n\nSpeaker bio: Ye Tian is a final-year Ph.D. student in Statistics at Columbia University. His research lies at the intersection of statistics\, data science\, and machine learning\, focusing on three main topics: (1) reliable transfer learning\; (2) high-dimensional statistics\; and (3) privacy and fairness of the learning system.
UID:130087-21865301@events.umich.edu
URL:https://events.umich.edu/event/130087
CLASS:PUBLIC
STATUS:CONFIRMED
CATEGORIES:seminar
LOCATION:West Hall - 411
CONTACT:
END:VEVENT
END:VCALENDAR