Implement a simple k-means method, Applied Statistics

Assignment Help:

There exists an unclassified data set with hidden data structures in it. The task in this assignment is to perform comprehensive Cluster Analysis in order to reveal the structures and similar data groups.

1. Implement a simple K-means method, which is able to handle real values data in attributes. Also you need to add functionality in your program that allows utilization of Euclidean, City Block, Euclidean Squared and Chebyshev distances. You are free to use any kind of weights (for feature or data instance) in the program if necessary.

2. Find unlabeled data set test.txt and initial centroids data set centroids.txt in the archive, both files have the following format: [attribute1_value attribute2_value ... attribute90_value]. The unlabeled data set includes 350 samples and the initial centroids set consists of 15 samples. Data instances in both files have 90 attributes.


Related Discussions:- Implement a simple k-means method

Determine the regression equation, The file Midterm  Data.xls has a tab lab...

The file Midterm  Data.xls has a tab labeled "National Grid vs. Alcoa" which presents historical price data for two stocks.  Using the National Grid price as the X-value and the Al

Data project, Dr. Jim Mirabella UNIT EIGHT: DATA ANALYSIS PROJECT All Excel...

Dr. Jim Mirabella UNIT EIGHT: DATA ANALYSIS PROJECT All Excel output should be copied into a single Word document where you must enter all of your responses to the questions below.

Econometrics, The following data on calcium content of wheat are consistent...

The following data on calcium content of wheat are consistent with summary quantities that appeared in the article “Mineral Contents of Cereal Grains as Affected by Storage and Ins

Multi stage or cluster random sampling, Multi stage or Cluster Random sampl...

Multi stage or Cluster Random sampling  Under this method, the random selection is made of primary, intermediate and final units from a given population. The area of investigat

Business reporting and analysis, You are a business analyst working for a c...

You are a business analyst working for a company called Combined Computers Pty Ltd. You have been asked to prepare a business report with statistics in it for the managing director

Box plot of income, The box plot displays the diversity of data for the inc...

The box plot displays the diversity of data for the income; the data ranges from 20 being the minimum value and 1110 being the maximum value. The box plot is positively skewed at 4

Sequential sampling, Sequential Sampling Under this method, a number of...

Sequential Sampling Under this method, a number of sample lots are drawn one after another from a universe depending on the results of the earlier samples. Such sampling is gen

Ryan-joiner - normal probability plot, The Null Hypothesis - H0:  The rando...

The Null Hypothesis - H0:  The random errors will be normally distributed The Alternative Hypothesis - H1:  The random errors are not normally distributed Reject H0: when P-v

What are the coefficients of the linear combination, For the following ques...

For the following questions we are interested in a comparison of the 16 years education vs. > 16 years. (Recall we did the analysis on the log scale, so these are actual means on t

Construct a cumulative percentage polygon, 1. For each of the following var...

1. For each of the following variables: major, graduate GPA, and height: a. Determine whether the variable is categorical or numerical. b. If the variable is numerical, deter

Write Your Message!

Captcha
Free Assignment Quote

Assured A++ Grade

Get guaranteed satisfaction & time on delivery in every assignment order you paid with us! We ensure premium quality solution document along with free turntin report!

All rights reserved! Copyrights ©2019-2020 ExpertsMind IT Educational Pvt Ltd