# AWS: MNIST and K-mean using sagemaker

**URL:** <https://community.cloudbank.org/t/aws-mnist-and-k-mean-using-sagemaker/178>\
**Category:** AWS\
**Created:** [June 24, 2024, 12:22am UTC](https://community.cloudbank.org/t/aws-mnist-and-k-mean-using-sagemaker/178 "2024-06-24T00:22:54Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![dchoi](https://avatars.discourse-cdn.com/v4/letter/d/85f322/32.png) [@dchoi](https://community.cloudbank.org/u/dchoi)\
**Post date:** [June 24, 2024, 12:22am UTC](https://community.cloudbank.org/t/aws-mnist-and-k-mean-using-sagemaker/178/1 "2024-06-24T00:22:54Z")

</div>

MNIST data set is applied for K-mean clustering using AWS sagemaker notebook.

This is an example procedure for a K-mean clustering using jupyter notebook on AWS sagemaker. I followed a tutorial shown in a github repository [1]

**Summary of the procedure**

- Download MNIST input data set and load it on S3.
- Load the input data and build K-mean model on sagemaker jupyter notebook.
- Get results of clustering classes and verifies. the model

**MNIST Input Data Loading**

- mnist.pkl.gz data file is found in the internet and downloaded. The file is directly uploaded to the notebook instance.
- using pickle module train and test data set are made.
- those input data is sent to the created S3 bucket object.

![image](https://us1.discourse-cdn.com/flex016/uploads/cloudbank1/original/1X/125d80216b73c2a50812ec1e20bfe3a8956ec812.png)

 ![image](https://us1.discourse-cdn.com/flex016/uploads/cloudbank1/original/1X/69f33e3c86adc2707be6c15c019468c4a0574630.png)  
 ![image](https://us1.discourse-cdn.com/flex016/uploads/cloudbank1/original/1X/1624c257756a311655f2d231ac2cfc774972ddf5.png)  
 ![image](https://us1.discourse-cdn.com/flex016/uploads/cloudbank1/original/1X/c0f23709f3e78d767edb796b998058072bab8701.png)

**K-Mean model build**  
K-mean model attributes are set and model is built using the input data. Here cluster number 10 is set for the number from 0 to 9. The model deployment is made.

![image](https://us1.discourse-cdn.com/flex016/uploads/cloudbank1/original/1X/3cdf7459cc49de1545b85844f6e86af53c49a37b.png)  
 ![image](https://us1.discourse-cdn.com/flex016/uploads/cloudbank1/original/1X/606a98ae2a7eccc8f6c7553950b81b2ecf156d8b.png)  
 ![image](https://us1.discourse-cdn.com/flex016/uploads/cloudbank1/original/1X/7db1b6c24f22327818f329f361dce333ddb94e01.png)

**Showing the output cluster**  
Top cluster classes willl be found. validation data is shown indeed belong to the classified cluster

 ![image](https://us1.discourse-cdn.com/flex016/uploads/cloudbank1/original/1X/3f551af3f78cf8e303140ef37ef3be7b5deeb2f9.png)  
 ![image](https://us1.discourse-cdn.com/flex016/uploads/cloudbank1/original/1X/cb250d3a8c7d52a932ea59a106880cebd176b899.png)  
 ![image](https://us1.discourse-cdn.com/flex016/uploads/cloudbank1/original/1X/01fed1ceff4f8cb84eccf0be627f16cef1548b5d.png)  
 ![image](https://us1.discourse-cdn.com/flex016/uploads/cloudbank1/original/1X/fe14847d6d46bcddf4ac01af6ae0006069957b9f.png)

Code used in this example is show in [4]

**Reference**  
[1] [GitHub - mtm12/SageMakerDemo](https://github.com/mtm12/SageMakerDemo)  
[2] [K-Means Algorithm - Amazon SageMaker](https://docs.aws.amazon.com/sagemaker/latest/dg/k-means.html)  
[3] [https://sagemakerexamples.readthedocs.io/en/latest/introduction\_to\_applying\_machine\_learning/US-census\_population\_segmentation\_PCA\_Kmeans/sagemaker-countycensusclustering.html](https://sagemakerexamples.readthedocs.io/en/latest/introduction_to_applying_machine_learning/US-census_population_segmentation_PCA_Kmeans/sagemaker-countycensusclustering.html)  
[4] [GitHub - dchoi/mnistKmean: mnistKmean](https://github.com/dchoi/mnistKmean)
