Praca magisterska podejmuje temat modelowania zapotrzebowania energetycznego budynków mieszkalnych w Polsce w kontekście ograniczeń dostępnych danych publicznych i ich wpływu na diagnozę ubóstwa energetycznego. Analiza wynika z potrzeby precyzyjnego szacowania kosztów energii w obliczu szoków cenowych z lat 2021–2023 oraz planowanych regulacji unijnych (EU-ETS2). Autor wskazuje na niewiarygodność oficjalnych danych GUS dotyczących struktury źródeł ciepła w gospodarstwach domowych.
Metodologia badań opiera się na integracji wielkoskalowych zbiorów danych: Centralnego Rejestru Charakterystyki Energetycznej (ponad 2,7 mln świadectw), Centralnej Ewidencji Emisyjności Budynków (CEEB) oraz danych demograficznych DataWise. Wykorzystując algorytmy uczenia maszynowego, w szczególności lasy losowe (Random Forest), opracowano modele regresyjne i klasyfikacyjne służące do estymacji wskaźnika zapotrzebowania na energię końcową dla budynków nieposiadających świadectw. Kluczowym elementem pracy jest podejście do mitygacji nadreprezentacji nowych budynków w zbiorze uczącym poprzez korektę ważącą strukturę meldunków oraz segmentację gmin. Celem pracy jest wykazanie potencjału zaawansowanej analityki danych w tworzeniu precyzyjnych polityk publicznych (evidence-based policy).
The master’s thesis addresses the modelling of residential building energy demand in Poland in the context of limited publicly available data and its impact on diagnosing energy poverty. The analysis stems from the need to accurately estimate energy costs in light of the 2021–2023 price shocks and planned EU regulations (EU ETS2). The author highlights the unreliability of official GUS data on the structure of heat sources in households.
The research methodology is based on integrating large-scale datasets: the Central Register of Energy Performance (over 2.7 million certificates), the Central Register of Building Emissions (CEEB), and DataWise demographic data. Using machine learning algorithms, particularly Random Forest, regression and classification models were developed to estimate the final energy demand index for buildings without certificates. A key element of the thesis is the approach to mitigating the overrepresentation of new buildings in the training set through weighting the report structure and segmenting municipalities. The aim of the thesis is to demonstrate the potential of advanced data analytics in developing precise, evidence-based public policies.




