One Month into Stanford Summer Session!

 Hello everyone!


I'm about a month into Stanford's Summer Session, where I am taking Principles of Data Science with Professor Alexander Dekhtyar and Introduction to Applied Statistics with Professor Madeline Louise Schroth-Glanz. So far, the experience has been incredible not only because of the amazing classes I'm taking but also because of the people I have met. Everyone here is so supportive and hardworking. I've never been among so many driven, intelligent people who, even in high school, have such specialized knowledge in their interests. When I speak to each person here, I can see how much learning means to them, and that pushes me to be an even better version of myself. 


Of course, the learners I have met here are amazing, but the classes are, too. In Principles of Data Science, we have begun learning machine learning algorithms that we will focus on for the remainder of the quarter. We are doing K-Nearest Neighbors now, but later we will learn about unsupervised learning algorithms like clustering. My favorite part of that class so far was learning about Simpson's paradox, especially when applying it to Titanic survivability. It turns out that although survivability was much higher for women than for men overall, it actually isn’t as clear as it seems because survival rates varied dramatically by passenger class, and men and women were not evenly distributed across those classes. Learning other classifiers like CountVectorizer and TF-IDF was also so cool; I never thought there was a quantifiable way to compare two similar pieces of text through metrics like cosine similarity. In one lecture, we uncovered which two Dr. Seuss books were the most similar by using TF-IDF, a vectorizer that assigns values to words and compares the overall words of one text to another. 


In my other class, Introduction to Applied Statistics, we have delved deeply into the regression line. While I have taken AP Statistics, this class was definitely a jump, especially since I had to build my beginner R skills before I could apply what I was learning. Most recently, we identified the difference between prediction and confidence intervals when predicting an NBA player's stats. In fact, predicting the range of a specific new NBA player's stats is much wider because the n value, or sample size, is smaller (in fact, it is just 1), making the standard error larger and thus the interval wider. A confidence interval is based on the mean of many players with the same value (e.g., minutes played), which decreases the standard error and thus the confidence interval. 

There is still almost a month left in my classes. We have reached the median lecture in both my classes, and I am excited to know I can learn as much as I have already learned again. It truly is an amazing experience and learning opportunity, and I can't wait to do my Principles of Data Science final project, where I can use what I've learned on something of my own.

This is just the beginning.

UPDATE: I have completed the program as of August 16th, 2026, and I am happy to say I received A's in both classes! I also wanted to share a link to my final project, which I completed with a classmate. It compares how we see NBA players vs. how they truly are as players (according to the stats), where the stats and eye test agree, and where they don't align.

Comments

Popular posts from this blog

My Response to a TED Talk by Mallory Freeman

Brookline Business Research | Part 2: The Data

Do the 3Cs (Cut, Color, Clarity) Really Impact the Overall Pricing of a Diamond?