---
type: work
area: work
status: active
date: 2026-05-03
created: 2026-05-03
updated: 2025-09-16
tags:
  - work
---
You are a senior data engineer with tons of years experience I want to help me resolve this problem and find the best solution. 
I have a project to create a database that contains a lot of data +336M row and 35 column just one table I will inject a data sample file of data to have an idea about it.
This data represent a power bi service dashboard extraction every month for every month i export all 12 month  before actual month and i store them in 12 parquet file.  
I have 42 extraction until august 2025 where in every extraction I extract 8 million row.
My main goal is to use this data in power bi mainly for analyzing the gap between 2 selected extraction, 
What type of database will serve me in that take in consideration that I will add monthly 8M row to this data?
My objectif is speed and performance in power bi so when I choose 2 extraction date from a slicer to compare them i want data to be filtered fast and data retrieval be also fast, How I can achieve that ?
Also I want to start with a local solution in my computer like I will start with 1à extraction and try all of that locally if it work as I except I will migrate to cloud 
Think deeply about this problem and answer my questions.
