Podcast

Data Warehouse Project Management: Challenges and Best Practices I.

Data Warehouse Project Management: Challenges and Best Practices I.

Tamás Fábián

4 min read

For a project manager following traditional project management methodologies, even a data warehouse project of average complexity can pose a massive challenge. [2]

In this article, we examine some aspects whose interpretation and application differ between traditional project management and DWH systems project management.

A DWH Anomaly: A successful project is one that never ends

A generalist project manager is usually accustomed to gathering requirements in the preparation phase, detailing them, discussing business implications, preparing IT specifications, and finally, after development and testing, delivering the product. 

However, one of the main characteristics of data warehouse projects is that often, a sufficiently deep detail of the business requirements is not available due to the lack of source system data or source system developments. At the same time, we know that some level of planning for these (ETL process, data modeling, DW-level storage, specifying DM-level logic) is essential; moreover, channeling data into business processes may involve additional data requirements that we must order even from the source systems. [3]

Co-existence with the vendor can also differ from what is experienced with other systems. For example, the requirements specification prepared for the development of a new module of a CRM system is relatively well-definable, so the implementation can be completed in a way that can be used for years without major (conceptual) modification.

In contrast, when we introduce a new data mart into a data warehouse, integrating a specific data domain often necessitates channeling additional data domains and source systems into the data warehouse. Employing vendors in such an environment on a project-based (fixed-price) basis is often only feasible at the expense of losing their sense of ownership, making it a less viable long-term solution. It is far more advantageous to establish a strategic partnership where the vendor (even if they do not legally own the source code) can implement developments and maintain data domains as a diligent custodian, while keeping the professional scope loose(r). 

Management-level positioning

In many cases, positioning data warehouse developments at the management level also poses a difficulty, as producing and loading a simple, data mart-level table can, in extreme cases, take man-months – which is quite a disappointing experience for, say, 10 fields.

 Anyone who has managed the organization and development of product usage or customer behavior metrics defined by data miners in a banking environment can easily see that carefully specifying a product usage probability based on linear regression, with moving variables and parameterizable by business users, is a time-consuming task. In extreme cases, this can take weeks, not to mention the development and extracting the variables from source systems.

"Garbage in, Garbage out", the importance of data quality

There is a data warehouse axiom: the strength of the DWH lies in its simplicity and robustness. However, it depends heavily on source systems, as it is highly sensitive to the accuracy, structure, and timely availability of the received data. If these systems do not operate with adequate integrity, or if the data they send cannot be considered reliable, this fundamentally defines the management-level perception of our data warehouse. 

"Garbage in, Garbage out" (GIGO) is often heard in English-speaking DWH circles, referring (also) to the incorrect business conclusions drawn from poor-quality data channeled into the data warehouse. 

As a manager, we must strive to navigate between the client business unit and the given source system when we need to channel data of questionable reliability into the data warehouse. It is an important aspect that loading and utilization in the data warehouse are well-specified, where basic business rules, value sets, and dimension data are specified in sufficient depth and quality, in a manner approved by the source system analysts. Reliable DM-level utilization can only be achieved this way. 

It is an important consideration that while we have no, or only very limited, influence on the data quality of a third-party provider's data (such as Opten, Cégtár, or Ministry of Interior data) when channeling it into the DW, as a data warehouse manager we can indeed maintain the demand to understand the quality and context of the internal data domains we wish to integrate. If we do not know what data assets we possess, we cannot expect the business units to know either.

In reality, business units using the DWH often focus on a single data domain (e.g., card data, company network, customer master, behavioral or product usage data, etc.) where they already have some level of information regarding the storage characteristics of the source system, or they already rely on these in their business processes. 

As a data warehouse manager, our goal is to ensure that data service is centralized, with appropriate data quality, and without bypassing the data warehouse. 

Scheduling batch processes

It is important to define in advance, as part of the analysis work, the frequency with which source system data is generated, the schedule and deadline by which the data warehouse can receive it, and what other data domains might still be required to produce a data mart-level report.

In the case of file transfer-based communication, the management of files arrived in the transfer area is typically performed by various custom scripts. Different batch processes are responsible for loading the data into Stage, DW, and then DM levels.

The latter are usually loading procedures run by scheduled jobs, organized so that the data enters the data warehouse as soon as possible, taking dependencies into account; all this in an automated manner, with minimal additional operational effort. 

The application used to manage the above is PWM, which can significantly shorten the lead time of Batch processes. Through self-learning algorithms built into the system, the resource consumption of DWH servers can be reduced, as well as the time spent on error handling.

The place for self-examination, DWH data corrections

In organizations where there are no high-quality data assets and yet they decide to implement a data warehouse, a situation can easily arise where, in the absence of a meaningful data governance strategy, only the infrastructure and experts are available, and no one takes ownership of the data. In such cases, management commitment is difficult to sustain in the long run, which can easily lead to the erosion of the data warehouse area. 

If we are forced to load data of questionable quality into our data warehouse, the need for correction will likely arise. Depending on the complexity and volume of the requirements, it is worth conducting a self-examination from time to time: 

  • Did the DWH analyst certainly understand the scope of the data to be channeled into the data warehouse in sufficient depth?

  • Did we sufficiently verify its data cleanliness?

  • Did the system analysts and managers on the source system side communicate honestly enough about the usability of the data?

  • Is there any circumstance we are unaware of? 

The question cannot be phrased completely generally. The point is that something went wrong, and this does not become acceptable just because it happens frequently. 

Data corrections and business parameterizations must also take place exclusively in an audited, historicized manner. The CLARA application, already well-known in the domestic banking and insurance sectors, can assist in this, as it can be implemented cost-effectively and quickly in a data warehouse environment. 

Regardless of this, it may be worth implementing corrections as early as possible on the source side, adhering to the general guideline that the data owner is always responsible for storing data in the appropriate quality. 

In this sense, the data warehouse will never be the data owner of the data stored at the DW level, but it will be for the DM-level data derived from them. 

Apart from these, there may be cases where we are forced to correct source-side, data quality-related errors on the data warehouse side, even if this is not the cleanest approach from a data governance perspective.

Related articles on the topic

Related articles on the topic

Budapest

1145 Budapest, Erzsébet Királyné útja 29/b.

Tel.: +36 1 422-3030

Fax: +36 1 422-3032

SZEGED

6724 Szeged, Bakay Nándor St. 24. Building D2

Tel.: +36 1 422-3030

Fax: +36 1 422-3032

Follow us on our platforms

© 2026 Clarity Consulting. All rights reserved.

Made by: ff. next

Széchenyi Plan 2020, European Union - Investing in your future support banner.
Széchenyi Plan 2020, European Union - Investing in your future support banner.
Széchenyi Plan 2020, European Union - Investing in your future support banner.

Budapest

1145 Budapest, Erzsébet Királyné útja 29/b.

Tel.: +36 1 422-3030

Fax: +36 1 422-3032

SZEGED

6724 Szeged, Bakay Nándor St. 24. Building D2

Tel.: +36 1 422-3030

Fax: +36 1 422-3032

Follow us on our platforms

© 2026 Clarity Consulting. All rights reserved.

Made by: ff. next

Széchenyi Plan 2020, European Union - Investing in your future support banner.
Széchenyi Plan 2020, European Union - Investing in your future support banner.
Széchenyi Plan 2020, European Union - Investing in your future support banner.

Budapest

1145 Budapest, Erzsébet Királyné útja 29/b.

Tel.: +36 1 422-3030

Fax: +36 1 422-3032

SZEGED

6724 Szeged, Bakay Nándor St. 24. Building D2

Tel.: +36 1 422-3030

Fax: +36 1 422-3032

Follow us on our platforms

© 2026 Clarity Consulting. All rights reserved.

Made by: ff. next

Széchenyi Plan 2020, European Union - Investing in your future support banner.
Széchenyi Plan 2020, European Union - Investing in your future support banner.
Széchenyi Plan 2020, European Union - Investing in your future support banner.