← Indietro
EsameEsame completoTesto d’esame

2016 04 27

Esame completo di Data Mining and Text Mining per il corso di Computer Engineering presso Politecnico di Milano. Materiale proveniente dall’archivio storico Studwiz e classificato per la consultazione online.

Data Mining and Text MiningEsame completo

Informazioni sul documento

Cosa trovi in questo materiale

Esame completo di Data Mining and Text Mining per il corso di Computer Engineering presso Politecnico di Milano. Materiale proveniente dall’archivio storico Studwiz e classificato per la consultazione online.

Qualità dell’importazione: il testo è stato estratto direttamente dal documento originale.

Contenuti estratti dal documento

Passaggi rappresentativi riconosciuti nelle diverse parti del materiale. Il testo completo resta presente nella pagina per la ricerca, mentre l’anteprima compatta rende più semplice la lettura.

Pagina 1

Politecnico di Milano School of Industrial and Information Engineering Data Mining and Text Mining Prof. Pier Luca Lanzi April 27, 2016 NAME CODICE PERSONA/ID • Solve the problems and write the answers inside the problem box. • Answers must be clearly written. • Pencils are not allowed. The midterm consists of 3 sheets of paper. It must be returned with all the 3 sheets. No any other sheet can be added. No sheet can be removed. • This is a closed-book/closed-notes exam. • Only non-programmable calculators are allowed. • Notes/books/mobile phones are not allowed. • If you are caught using forbidden material, the exam will immediately end and an RP grade will be recorded; then, your Data Mining exam will consist of an oral examination from then on. • All the answers must be adequately motivated. Grades Problem 1. (7 points) Consider a set of six data points (A, B, C, D, E and F) and the following distance matrix: A B C D E F A | 0 | B | 2 0 | C | 1 1 0 | D | 7 5 6 0 | E | 8 6 7 1 0 | F | 5 3 4 2 3 0 | 1. Apply hierarchical clustering using the complete link (or MAX) approach to measure distance between clusters and show the final dendrogram. 2. Suppose that instead of the complete link, we use the centroid distance to compute the distance between clusters, would it be correct to say that then hierarchical clustering would become equivalent to k-means? Problem 1. (continued) Problem 2. (7 points) (A) Compute the FP-Tree for the following transactions using a minimum support count of 4. (B) What is lift in the context of association rule mining? What does it represent? A B C D E F G H I 0 1 1 1 0 1 1 0 0 1 1 1 1 1 1 0 0 0 1 0 0 1 1 1 0 1 0 1 1 1 1 0 0 1 0 0 1 0 1 1 1 1 0 0 0 1 1 1 0 1 0 0 0 0 1 1 1 0 0 0 1 0 0 0 1 0 0 0 0 1 0 1 Bonus Question (2 points). Given the FP-Tree…

Anteprima

Prima pagina del documento.

Prima pagina: 2016 04 27