Informazioni sul documento
- Università
- Politecnico di Milano
- Corso di laurea
- Computer Engineering
- Materia
- Data Mining and Text Mining
- Classificazione
- Esame · Esame completo
- Contenuto
- Testo d’esame
- Formato originale
- Testo
- Testo ricercabile
Esame completo di Data Mining and Text Mining per il corso di Computer Engineering presso Politecnico di Milano. Materiale proveniente dall’archivio storico Studwiz e classificato per la consultazione online.
Esame completo di Data Mining and Text Mining per il corso di Computer Engineering presso Politecnico di Milano. Materiale proveniente dall’archivio storico Studwiz e classificato per la consultazione online.
Qualità dell’importazione: il testo è stato estratto direttamente dal documento originale.
Passaggi rappresentativi riconosciuti nelle diverse parti del materiale. Il testo completo resta presente nella pagina per la ricerca, mentre l’anteprima compatta rende più semplice la lettura.
Politecnico di Milano School of Industrial and Information Engineering Data Mining and Text Mining Prof. Pier Luca Lanzi & Ing. Daniele Loiacono November 23, 2016 NAME CODICE PERSONA/ID • Answers must be clearly written inside the problem box. All the answers must be adequately motivated. • Pencils are not allowed. The midterm consists of 5 sheets of paper. It must be returned with all the 5 sheets. No any other sheet can be added. No sheet can be removed. • This is a closed-book/closed-notes exam. • Only non-programmable calculators are allowed. • Notes/books/mobile phones are not allowed. • If you are caught using forbidden material, the exam will immediately end and an RP grade will be recorded; then, your Data Mining exam will consist of an oral examination from then on. • Scoring o A problem left unsolved will amount to zero points. o A completely wrong solution will amount to -3 points Grades Problem 1 (7pts). Compute the frequent itemsets using FP-Growth using a support of 0.4 Problem 1. (continued) Problem 2 (7pts). Briefly define the concept of completeness and optimization as presented during the course and discuss them in the context of decision trees and association rules. Problem 3 (5pts). The company ZoolanderData asked three companies DataOne, ShouldClassify, and ProbData to help them compare four algorithms a basic decision tree returning class labels, a naïve bayes, Logistic regression, and a version of k-nn modified to return probabilities values. DataOne compared the four algorithms using ROC curves. The result shows that the decision tree is the best performing algorithm with the largest area below the curve. ShouldClassify applied crossvalidation and compared the four algorithms. Their results show confirm the results presented by DataOne. ProbData…
Prima pagina del documento.