ch
Feedback
Data Engineering / Инженерия данных / Data Engineer / DWH

Data Engineering / Инженерия данных / Data Engineer / DWH

前往频道在 Telegram

Data Engineering: ETL / DWH / Data Pipelines based on Open-Source software. Инженерия данных. ✔ DWH / SQL ✔ Airflow / Python / ETL / dbt / Spark ✔ AI Agents Рекламу не размещаю Вопросы: @iv_shamaev | datatalks.ru

显示更多
2 690
订阅者
+124 小时
+57
+2130
帖子存档
GitHub - ClickHouse/clickhouse-presentations: Presentations, meetups and talks about ClickHouse https://github.com/ClickHouse/clickhouse-presentations

Automate without limits n8n The workflow automation platform that doesn't box you in, that you never outgrow GitHub 27k+ Usage 🔹 Learn how to install and use it from the command line 🔹 Learn how to run n8n in Docker Self-Hosted -> Free 🔹 Data stays on your infrastructure 🔹 Open & extendable 🔹 One-line npm command or Docker deployment Habr: n8n. Автоматизация ИБ со вкусом смузи

Open Source Guides Open source software is made by people just like you. Learn how to launch and grow your project. https://opensource.guide/

Mara Pipelines This package contains a lightweight data transformation framework with a focus on transparency and complexity reduction. It has a number of baked-in assumptions/ principles: - Data integration pipelines as code: pipelines, tasks and commands are created using declarative Python code. - PostgreSQL as a data processing engine. - Extensive web ui. The web browser as the main tool for inspecting, running and debugging pipelines. - GNU make semantics. Nodes depend on the completion of upstream nodes. No data dependencies or data flows. - No in-app data processing: command line tools as the main tool for interacting with databases and data. - Single machine pipeline execution based on Python's multiprocessing. No need for distributed task queues. Easy debugging and output logging. - Cost based priority queues: nodes with higher cost (based on recorded run times) are run first. https://github.com/mara/mara-pipelines

Глубокое погружение в Data Quality / Хабр https://habr.com/ru/company/vk/blog/674876/

Проектирование ETL-пайплайна в Apache Airflow / Хабр https://habr.com/ru/company/otus/blog/679402/

GitHub - martandsingh/ApacheSpark: This repository will help you to learn about databricks concept with the help of examples. It will include all the important topics which we need in our real life experience as a data engineer. We will be using pyspark & sparksql for the development. At the end of the course we also cover few case studies. https://github.com/martandsingh/ApacheSpark

ETL Pipeline with Airflow, Spark, s3, MongoDB and Amazon Redshift Educational project on how to build an ETL (Extract, Transform, Load) data pipeline, orchestrated with Airflow. https://github.com/renatootescu/ETL-pipeline

🔥 Awesome Docker Compose samples These samples provide a starting point for how to integrate different services using a Compose file and to manage their deployment with Docker Compose. 👉 @devops_dataops https://github.com/docker/awesome-compose

Software Engineering for Absolute Beginners - 2021 What You Will Learn 🔹 Explore the concepts that you will encounter in the majority of companies doing software development 🔹 Create readable code that is neat as well as well-designed 🔹 Build code that is source controlled, containerized, and deployable 🔹 Secure your codebase 🔹 Optimize your workspace

Как собрать платформу обработки данных «своими руками»? @devops_dataops https://habr.com/ru/company/itsumma/blog/679516/

Продвинутая работа с Docker — Docker-compose. https://1cloud.ru/blog/docker-compose