Posts

Filler Journal - 10/18/19 - 11/15/19

Image
     On Thursday, November 14th, I ran into a multitude of issues. All of which was because of pent up anxiety and downright disappointment. It all began in my debate class. I had a tournament coming up soon and my case was barely finished. I know I could finish it late in the day but I felt the need to do it during class. I was compelled to do so since I wanted to do well but I ended up not because I had a math test one period away. For the rest of the class, I did a last-minute cram session but I guess that was bad luck. Right before the test, my teacher, Mrs. Linton, said that if we had 6 minutes left, we "should get on the calculator section ASAP." Judging from her statement, I wrongfully judged that I would need around 10 minutes since I thought she implied that the calc section would take 6. Instead, I really needed much more time. I ended up doing extremely poor and got disappointed. During my "alone" time of thinking about what could I have done to do better...

Zachary's Karate Club Technical Journal 3

Image
General Overview The Girvan-Newman Algorithm, named after Michelle Girvan and Mark Newman, is a hierarchical method used to detect communities in a graph depending on the iterative elimination of edges with the highest number of shortest paths that go through them. Essentially the algorithm is the process of repetitive calculations of edge betweenness centrality of all nodes and cutting the edge with the highest centrality. In pseudo-code: R epeat until no edges are left: C alculate edge betweenness for every node in the graph R emove the edge with the highest edge betweenness R ecalculate edge betweenness for all remaining nodes C onnected nodes are considered communities This method works because groups are separated from one another, revealing the underlying community structure of the network. Edge betweenness is the driving factor of the algorithm but it must be recalculated after each completed iteration. In sentences, edge betweenness works like the following: for ev...

Zachary's Karate Club Technical Journal 2

Image
The Problem:   My last technical journal for this dataset can be found on the home page but here's a short summary: Zachary’s Karate Club is a social network of a karate club that captures 34 members and records links with pairs of interacting members as edges. During the study, the karate club had split into two (between John and Mr. Hi), coincidentally also dividing the members into the two factions. The task at hand is to determine/predict which members would stick with John or leave with Mr. Hi. This is an example of an algorithm splitting the 34 members into two groups, indicated by the color difference The Algorithm: Previously I've tried many algorithms like K-nearest-neighbor but all of those have proven not applicable to this graph dataset. However , I was fortunate enough to have found an algorithm that sounds promising: Community Detection/Girvan-Newman Algorithm. Here's a basic rundown on how the code works: Girvan-Newman Algorithm (In...

Graph Clustering Algorithms - 10/11-25/19

Image
On Thursday, we went to Caltech to visit Dr. Hassibi for our biweekly meeting. Each of us would talk about our algorithm. Dr. Hassibi went over Spectral Clustering which works like finding the max eigenvalue and its eigenvector to create a k-number of cliques. Cliques are the ideal graph clusters because all the nodes in a cluster are connected to each other and never to other nodes outside. Our task for the next three weeks is to complete testing for our algorithms on the Karate Club dataset and understand the math behind it. Picture of Dr. Hassibi and the whiteboard detailing the basics of what Spectral clustering is I also went to West Torrance High School to do a math competition (BML) with my friends. I took the Pre-calculus, Calculus, and Number Theory tests, all of which were a challenge to me. I had studied for an extensive amount of time but my scores did not reflect my efforts. I tripped up on careless mistakes and some unexcusable mishaps ...

Filler Journal - 9/30-10/27/19

Image
Yesterday I already wrote about a Technical Journal detailing what I did for the last two weeks in my Caltech STEM program. I went over the Zachary's Karate Club dataset and predicting the members with Kmeans and Spectral Clustering Algorithm. So in substitute for such, I'll just talk about what has been on my mind for the past few weeks. My APUSH class has not been that stressful but the tests do worry me a lot. Mr. Paccone, my teacher, has recently assigned us large parts of his lesson slideshow and tested us given a short time. And by the time we take the test, I have absorbed nothing from cram sessions with the slides. Thankfully I watched videos from the internet to cope with my lack of understanding. My English class has been fine with me. We're reading this book called The Crucible  which is based on the Salem Witch Trials held in Massachusetts. The purpose of the novel was to address McCarthyism and its obvious problems. I've taken a quiz and test on it...

Zachary's Karate Club Technical Journal 1

Image
The Problem: The data set we needed to correctly predict was Zachary’s Karate Club, a social network of a university karate club studied by Wayne W. Zachary between 1970 and 1972. The network captures 34 members of a karate club and records links with pairs of interacting members. During the study, the owner of the club, John A, and instructor, Mr. Hi, had an argument and decided to part ways from each other. This effectively split the 34 members into two: half of them left for Mr. Hi’s new company and the rest either found a new instructor or quit the sport. The task at hand is from a data set, determine which members would stick with John or leave with Mr. Hi. This is an example of an algorithm splitting the 34 members into two groups, indicated by the color difference The Data: The data, in its simplest form, is a graph that has 34 nodes, each one representing one member. The nodes have a node attribute ‘club’ that indicates the name of the club to which the member ...

Clustering Algorithms - 9/16-27/19

Image
The past two weeks have been more towards our field of study than what happened previously. Prior to the second meeting with Dr. Hassibi, the Machine Learning group was given the task to learn about Cluster Algorithms. We spent time studying algorithms such as Expectation-Maximization, k-means and Gaussian Mixture Model. The objective with each subtopic was to understand its purpose, what makes it different from other methods, and the math behind it. This was much more ambiguous than learning what it does. The learning process was somewhat very uncoordinated because situations where you needed to learn an idea to understand a particular thing, leading to more and more topics for us to pick up on. Once all algorithms have been inferred thoroughly, each student was given a topic to write a concept map on. Below is an extensive video going over what Clustering is in Machine Learning: I was assigned to the Expectation-Maximization Algorithm which is an iterative method to ...