Data Engineering / Инженерия данных / Data Engineer / DWH
الذهاب إلى القناة على Telegram
Data Engineering: ETL / DWH / Data Pipelines based on Open-Source software. Инженерия данных. ✔ DWH / SQL ✔ Airflow / Python / ETL / dbt / Spark ✔ AI Agents Рекламу не размещаю Вопросы: @iv_shamaev | datatalks.ru
إظهار المزيد2 694
المشتركون
+124 ساعات
+57 أيام
+2130 أيام
أرشيف المشاركات
Репозиторий с проектами Data Engineering
https://github.com/san089/Udacity-Data-Engineering-Projects
GitHub - ClickHouse/clickhouse-presentations: Presentations, meetups and talks about ClickHouse
https://github.com/ClickHouse/clickhouse-presentations
Automate without limits n8n
The workflow automation platform that doesn't box you in, that you never outgrow
GitHub 27k+
Usage
🔹 Learn how to install and use it from the command line
🔹 Learn how to run n8n in Docker
Self-Hosted -> Free
🔹 Data stays on your infrastructure
🔹 Open & extendable
🔹 One-line npm command or Docker deployment
Habr: n8n. Автоматизация ИБ со вкусом смузи
Open Source Guides
Open source software is made by people just like you. Learn how to launch and grow your project.
https://opensource.guide/
Mara Pipelines
This package contains a lightweight data transformation framework with a focus on transparency and complexity reduction. It has a number of baked-in assumptions/ principles:
- Data integration pipelines as code: pipelines, tasks and commands are created using declarative Python code.
- PostgreSQL as a data processing engine.
- Extensive web ui. The web browser as the main tool for inspecting, running and debugging pipelines.
- GNU make semantics. Nodes depend on the completion of upstream nodes. No data dependencies or data flows.
- No in-app data processing: command line tools as the main tool for interacting with databases and data.
- Single machine pipeline execution based on Python's multiprocessing. No need for distributed task queues. Easy debugging and output logging.
- Cost based priority queues: nodes with higher cost (based on recorded run times) are run first.
https://github.com/mara/mara-pipelines
Примерчик ETL pipeline на python
https://github.com/iamaziz/etl
Глубокое погружение в Data Quality / Хабр
https://habr.com/ru/company/vk/blog/674876/
Проектирование ETL-пайплайна в Apache Airflow / Хабр
https://habr.com/ru/company/otus/blog/679402/
GitHub - martandsingh/ApacheSpark: This repository will help you to learn about databricks concept with the help of examples. It will include all the important topics which we need in our real life experience as a data engineer. We will be using pyspark & sparksql for the development. At the end of the course we also cover few case studies.
https://github.com/martandsingh/ApacheSpark
ETL Pipeline with Airflow, Spark, s3, MongoDB and Amazon Redshift
Educational project on how to build an ETL (Extract, Transform, Load) data pipeline, orchestrated with Airflow.
https://github.com/renatootescu/ETL-pipeline
🔥 Awesome Docker Compose samples
These samples provide a starting point for how to integrate different services using a Compose file and to manage their deployment with Docker Compose.
👉 @devops_dataops
https://github.com/docker/awesome-compose
Software Engineering for Absolute Beginners - 2021
What You Will Learn
🔹 Explore the concepts that you will encounter in the majority of companies doing software development
🔹 Create readable code that is neat as well as well-designed
🔹 Build code that is source controlled, containerized, and deployable
🔹 Secure your codebase
🔹 Optimize your workspace
Как собрать платформу обработки данных «своими руками»?
@devops_dataops
https://habr.com/ru/company/itsumma/blog/679516/
Список полезных Linux Commands
@devops_dataops
https://telegra.ph/Linux-Commands-07-27
Продвинутая работа с Docker — Docker-compose.
https://1cloud.ru/blog/docker-compose
Завтра в 12 трансляция
https://youtu.be/jF3YemOVofQ
