← Indietro
AppuntiCompleti

Summary notes of the course

Completi di Technologies for Information Systems per il corso di Computer Engineering presso Politecnico di Milano. Materiale proveniente dall’archivio storico Studwiz e classificato per la consultazione online.

Technologies for Information SystemsCompleti

Informazioni sul documento

Cosa trovi in questo materiale

Completi di Technologies for Information Systems per il corso di Computer Engineering presso Politecnico di Milano. Materiale proveniente dall’archivio storico Studwiz e classificato per la consultazione online.

Qualità dell’importazione: il testo è stato estratto direttamente dal documento originale.

Contenuti estratti dal documento

Passaggi rappresentativi riconosciuti nelle diverse parti del materiale. Il testo completo resta presente nella pagina per la ricerca, mentre l’anteprima compatta rende più semplice la lettura.

Pagina 1

Data Integration Introduction Data Integration: is the problem of combining data coming from different data sources, providing the user with a unified vision of the data, detecting correspondences between similar concepts that come from different sources, and conflicting solving. The aim of Data Integration is to set up a system where it is possible to query different data sources as if they were a unique one (through a global schema). Data integration is needed because of a need for interoperability among SW applications, services and information managed by different organizations: - find information and processing tools, when they are needed, independently of physical location - understand and employ the discovered information and tools, no matter what platform supports them, whether local or remote - evolve a processing environment for commercial use without being constrained to a single vendor’s offerings. Data heterogeneities: - same data model, different systems —> technological heterogeneity - different data models OR semi- or unstructured data (HTML, XML, multimedia..)—> model heterogeneity - same data model, different query languages —> language heterogeneity The four V’s of Big Data: - Volume: each data source contains a huge volume of data, and the number of data sources has grown. - Velocity: data is continuously made available and many of the data sources are very dynamic. - Variety: data sources are extremely heterogeneous both at the schema level, regarding how they structure their data, and at the instance level, regarding how they describe the same real world entity. - Veracity: is the degree to which data is accurate, precise and trusted. STEPS OF DATA INTEGRATION 1) Schema reconciliation (if the sources have a schema): mapping the data structure as in the…

Anteprima

Prima pagina del documento.

Prima pagina: Summary notes of the course