Back
NotesComplete set

Completed notes of the course

Complete course materials for Model Identification and Machine Learning in the Biomedical Engineering degree programme at Politecnico di Milano. The document covers: Machine learning Lecture 1: Data prepara-on Business intelligence systems and mathema1cal models for decision making can achieve accurate and effec1ve results only when the input data are highly reliable. However, the data extracted from the available primary sources may have

Model Identification and Machine LearningComplete set

Document information

What's included in this study material

Complete course materials for Model Identification and Machine Learning in the Biomedical Engineering degree programme at Politecnico di Milano. The document covers: Machine learning Lecture 1: Data prepara-on Business intelligence systems and mathema1cal models for decision making can achieve accurate and effec1ve results only when the input data are highly reliable. However, the data extracted from the available primary sources may have

Import quality: text was extracted directly from the original document.

Extracted content from the document

Representative passages recognised in different parts of the material. The full extracted text remains available to search, while this compact preview makes the page easier to read.

Page 1

Machine learning Lecture 1: Data prepara-on Business intelligence systems and mathema1cal models for decision making can achieve accurate and effec1ve results only when the input data are highly reliable. However, the data extracted from the available primary sources may have several anomalies which analysts must iden1fy and correct. Several techniques are employed to reach this goal. Data valida)on: the quality of input data may prove unsa1sfactory due to: • Incompleteness: some records may contain missing values corresponding to one or more aCributes. Data may be missing because of malfunc1oning recording devices. It is also possible that some data were deliberately removed during previous stages of the gathering process. Incompleteness may also derive from a failure to transfer data from the opera1onal databases to a data mart used for a specific business analysis or simply to the fact that that data was not mandatory. • Noise: data may contain erroneous or anomalous values, which are usually referred to as “outliers”. Other possible causes of noise are to be sought in malfunc1oning devices for data measurement, recording and transmission. • Inconsistency: some1mes data contain discrepancies due to changes in the coding system used for their representa1on, and therefore may appear inconsistent. The purpose of data valida1on techniques is to iden1fy and implement correc1ve ac1ons in case of incomplete and inconsistence data or data affected by noise. Incomplete data To par1ally correct incomplete data, there are several techniques: • Elimina1on: it is possible to discard all records for which the values of one or more aCributes are missing. This is done if at least 10% of the data is missing, we can eliminate a whole column or row. This is a dras1c and not thorough…

Preview

First page of the document.

First page: Completed notes of the course