DSC 232R: Big Data Analytics Using Spark
Jan 5, 2026
·
1 min read
Course Overview
DSC 232R: Big Data Analytics Using Spark is a graduate-level course in the Master of Data Science program at UC San Diego’s Halicioglu Data Science Institute.
Topics Covered
- Distributed computing fundamentals
- Apache Spark architecture and programming model
- Large-scale data processing with PySpark
- Machine learning at scale with MLlib
- Real-world applications in genomics and industry
- Cloud computing and cluster management
Learning Outcomes
Students completing this course will be able to:
- Design and implement distributed data processing pipelines
- Apply machine learning algorithms to large-scale datasets
- Optimize Spark applications for performance
- Work with real-world big data from genomics and other domains
Offering Schedule
- Spring 2024
- Fall 2025
- Winter 2026

Authors
Edwin Solares
(he/him)
Executive Director, ESB AI Lab Corporation
Executive Director of ESB AI Lab Corporation, a 501(c)(3) nonprofit advancing
research in AI, machine learning, computer vision, and genomics. Previously a
Lecturer at UC San Diego. My research harnesses AI and bioinformatics for food
security and species conservation. Published in Nature Plants, PNAS, Genome
Research, and G3 (h-index: 7).