Thursday, November 19, 2009

ETL

ETL - extract, transform, load

ETL architecture is used in order to manage and optimize the data together with the process of its transfer which, as a result, leads to simplification and standardization of data storage in the data warehouse.

Purpose of the ETL tool is to create a universal framework for different processes of transforming data into target destination.

The ETL letters are the abbreviations for extraction, transformation and loading of data, which means that the ETL tools extract data from a source, transform it into new formats and load it into target data warehouse.

System’s architects created a scheme which eliminates repeatedly appearing actions. Loading the data into warehouse requires two types of actions: one, in which data is transformed from the identifiable source system according to source-specific rules – this action results in transforming data into a standardized format. Each action has to be done by individual method for every data storage system.

The other, conformed type of ETL architecture means creating universal way of data entity, regardless of the source and purpose; it allows to follow reusable rules which are applied in business and work in different conditions with variety of data. Both systems has pros and cons. At the beginning the first one requires fewer actions to be performed in order to work smoothly and seems easier. However, looking at the overall life cycle of the data warehouse, in the long run using the conformed ETL is more efficient.

It is because data warehouse is being developed and has been growing continually when the new data is added to the system, so when the traditional architecture is taken into consideration, though it may seem easier to make it, in the long run it turns out that the same or similar actions are repeated, needlessly lengthening the development time and increasing its complexity. Implementing the universal for one entity ELT architecture requires more work than doing the same for a particular task, however in the long run it brings more benefits as the data is stored and the scheme may be used further.

Thus the most obvious advantage of using ETL is that the storage data is easily reused which results in improving its quality. Moreover it’s easier to make and use, because in fact it is less complex than storing data in different traditional systems. Additional advantage is also that it’s less difficult to add and acquire new source systems. One should also bear in mind that the conform ETL architecture is not to be applied everywhere, it should be proceeded by detailed analysis, which would ensure the benefits that would possibly result from following the conform architecture. Otherwise it may not be so beneficial.

This data architecture enables direct access to data in operational systems, which is very important and in fact shorten the time considerably, which otherwise would be spent on data implementation and simplify the data usage.

EII – Enterprise Information Integration

EII – Enterprise Information Integration

This fine abbreviation applies to information integration only at the commercial line. EII is a data integration from the many systems to compact and unified form that offered viewing and manipulation of data. Data is matching, selected and restructured and only then present to the user.
Data integration is primarily connected with manipulation and the valuation of historical data to discover existing trends which for example are not visible.
On the other hand, application integration is focused on data integration between application library and system. When data is changing in one system, the reversal is instituted to the other system of interest, usually via irregular messaging. Information integration is simply target at the end users who are working with multiple systems.
As You can see and as stated earlier, data integration is focused on the integration of data in multiple systems which I called “big tree”. In fact, not many people know about this database in spite of maybe accidental

Data Integration

Have You ever heard about Data Integration?
Do You know what exactly it is?

Data Integration - it is almost as simple as it sounds and it is pretty much intangible goal which software industry is still running for. In this article I will try to explain why data integration is so huge bite for this companies.

Data integration is like big tree where every branch connected and cooperating with the other. All the data living in this “branch” affords to the users unification view of this data. The process of data integration occurring in every part of our live. For example: when two similar company have to merge their databases or at education to merge different results. However, in our contention world the ITC infrastructure played major role in a learning environment which can drawn in the right students, sponsoring and research project. The most important think is the role to data integration. It means that two system can be integrate when:
Both look the same;
Act seems to be the same;
and both of this system produce and consume the same data.
In sum, data integration is the ability to control, make, share and consume the same data.

Example

Some web application have information about the cities such a weather, demographics, cinemas, hotels, tourism, etc. Typically, the information have to be in one type of database with a single schema. But it is not as simple as seems to be. Single enterprise would find the information which can be difficult to collect. Even if resources gather volumes of information about weather or tourism, it would probable duplicate the data.

The originators of this idea try to develop the best model of virtual schema that the users want the most. This wrapper escorted to resources or application to improve on compatibility of the crime database weather websites. This adapters for each data source transform the query results into a straight form. Finally, the database connect the performance to show unified view and answer for the users question.