← Indietro
EsameEsame completoTesto d’esame

2014 07 14

Esame completo di Data Mining and Text Mining per il corso di Computer Engineering presso Politecnico di Milano. Materiale proveniente dall’archivio storico Studwiz e classificato per la consultazione online.

Data Mining and Text MiningEsame completo

Informazioni sul documento

Cosa trovi in questo materiale

Esame completo di Data Mining and Text Mining per il corso di Computer Engineering presso Politecnico di Milano. Materiale proveniente dall’archivio storico Studwiz e classificato per la consultazione online.

Qualità dell’importazione: il testo è stato estratto direttamente dal documento originale.

Contenuti estratti dal documento

Passaggi rappresentativi riconosciuti nelle diverse parti del materiale. Il testo completo resta presente nella pagina per la ricerca, mentre l’anteprima compatta rende più semplice la lettura.

Pagina 1

Politecnico di Milano Facoltà di Ingegneria dell’Informazione Data Mining and Text Mining Tecniche di Apprendimento Automatico Prof. Pier Luca Lanzi & Ing. Daniele Loiacono July 14, 2014 NAME MATRICOLA Solve the following problems and write the answer inside the problem box. Answers must be clearly written. Pencils are not allowed. The midterm consists of 3 sheets of paper. It must be returned with all the 3 sheets. No any other sheet can be added. No sheet can be removed. This is a closed-book, closed-notes exam. Only non-programmable calculators are allowed. Notes/books/mobile phones are not allowed. If you are caught using forbidden material, the exam will immediately end and an RP grade will be recorded; then, your Data Mining exam will consist of an oral examination from then on. All the answers must be adequately motivated. Grades The image cannot be displayed. Your computer may not have enough memory to open the image, or the image may have been corrupted. Restart your computer, and then open the file again. If the red x still appears, you may have to delete the image and then insert it again. Problem 1. Compute the estimated accuracy of k-nearest neighbor using k = 1, Euclidean distance, and a tenfold cross validation on the following dataset where X and Y are the attributes and C is the class. X Y C 1 2 7 1 2 10 3 1 3 1 7 0 4 5 8 0 5 10 1 0 6 3 5 0 7 3 6 0 8 2 2 1 9 4 9 0 10 6 9 1 Problem 1. (continued) Problem 2. Consider the following dataset in which three clusters can be clearly seen. Apply three iterations of k-means starting with the centroids at (4,10), (2.5,2) and (10,5) using Euclidean distance. x y 1 -0.9 8.5 2 -2.0 7.7 3 -0.3 8.2 4 -0.3 10.0 5 8.4 3.0 6 8.9 1.7 7 9.4 1.0 8 8.2 1.7 9 -0.3 -0.2 10 0.2 0.1 11 1.0 0.1 12 0.8 0.4 13 -0.7 8.7 14 1.2 10.1…

Anteprima

Prima pagina del documento.

Prima pagina: 2014 07 14