1-day current streak·30-day longest streak
👋 Well Hello There! ✨ I’m Selenge Tulga , I’m a Data Engineer with 7 years (4.5 DE + 2.5 SWE) of experience and a background in Computer Science (B.S.)…
👋 Well Hello There! ✨
I’m Selenge Tulga , I’m a Data Engineer with 7 years (4.5 DE + 2.5 SWE) of experience and a background in Computer Science (B.S.) and Data Science (M.S.). I focus on building reliable, production-grade data platforms that replace manual reporting and support finance, operations, and business decision-making.
I enjoy working at the intersection of data systems, analytics, and real-world constraints, with a strong emphasis on data correctness, observability, and long-term maintainability. I care less about flashy tools and more about systems that are easy to reason about and trust.
📍 Based in Austin, TX 🇺🇸


---
🛠️ Core Stack
Data Engineering & Analytics
- Warehousing & Modeling: Snowflake, dbt, PostgreSQL, Redshift
- Pipelines & Orchestration: Airflow, Prefect, AWS Glue, ETL/ELT
- Streaming & Processing: Kafka, Flink, Spark, Databricks (when scale or latency requires it)
- Cloud: AWS (primary), GCP
- Visualization: Tableau, Power BI, Amazon QuickSight
---
🎓 Certifications
- AWS Certified Data Engineer – Associate
- AWS Certified Machine Learning Engineer – Associate
- AWS Certified Machine Learning – Specialty
- Databricks Certified Data Engineer – Associate
- Databricks Certified Data Engineer – Professional
🚀 What I’m Focused On Now
- Real-time trending topics pipeline: Kafka → PySpark Structured Streaming → AWS lakehouse (Parquet, idempotent writes)
- AI-powered data quality monitor: Airflow + dbt + Great Expectations + Claude API → Slack alerts in under 30 seconds
- Practicing failure handling, monitoring, and data quality patterns
-
realtime_reddit_trend ★ PINNED
End-to-end streaming data pipeline that ingests Reddit posts via the PRAW API, processes them with PySpark Structured Streaming, stores results in a serverless lakehouse (S3 + Glue + Athena), and visualises live trends in a Streamlit dashboard.
Python ★ 0 3mo agoExplain → -
ai_quality_check ★ PINNED
An end-to-end data quality pipeline that uses Claude AI to automatically diagnose root causes when data quality checks fail — and delivers plain-English incident reports to Slack and a live Streamlit dashboard. Built on Apache Airflow, dbt, DuckDB/MotherDuck, and Great Expectations
Python ★ 1 3mo agoExplain → -
order_tracking_power_bi ★ PINNED
This Power BI project provides insights into customer orders and product tracking using interactive dashboards. It visualizes order status, sales trends, shipping details, and warehouse assignments.
★ 10 11mo agoExplain → -
HR-Analysis ★ PINNED
A Power BI dashboard analyzing HR data, visualizing employee distribution, satisfaction, training costs, and department-wise insights.
★ 4 1y agoExplain → -
DE_AI_Interviewer ★ PINNED
An AI-powered mock interview tool that reads a real job description and generates 12 tailored interview questions across 6 data engineering pillars. Answer everything at once, then get a full scored debrief instantly. Built with React + Vite on the frontend and Express + Claude API on the backend.
JavaScript ★ 13 2mo agoExplain → -
interview_random_question_picker ★ PINNED
A Streamlit app that randomly picks interview practice questions by category.
Python ★ 4 2mo agoExplain → -
DSA
This repository is dedicated to solving problems from platforms like LeetCode and NeetCode with the goal of mastering DSA concepts and improving problem-solving skills.
Python ★ 2 10mo agoExplain → -
system-design-masterclass ⑂
This material is based on the System Design Masterclass (2025) course available on Udemy. You can find the course here: System Design Masterclass. You can also find many free resources related to this on Sweet Codey Official Website.
★ 2 1y agoExplain → -
selengetu.github.io ⑂
My self coded personal website build with React.js
JavaScript ★ 2 1y agoExplain → -
awesome-public-datasets ⑂
A topic-centric list of HQ open datasets.
★ 1 1y agoExplain → -
neetcode-submissions-gzmz9xhb
My NeetCode.io problem submissions
Python ★ 0 1d agoExplain → -
selengetu
No description.
★ 0 2mo agoExplain → -
non-coding-skills ⑂
Claude skills for personal (non-coding) workflows.
★ 0 3mo agoExplain → -
Python_scripts
No description.
Python ★ 0 3mo agoExplain → -
Nba
End-to-end NBA data engineering pipeline built with Python, Apache Airflow, dbt and Snowflake. The project ingests data from public NBA APIs, enforces data quality checks, logs ingestion metadata for observability, and loads dimensional and fact tables into a cloud data warehouse, following production-style data engineering best practices.
Python ★ 0 3mo agoExplain → -
ThinkPython ⑂
Jupyter notebooks and other resources for Think Python by Allen Downey, published by O'Reilly Media.
★ 0 6mo agoExplain → -
abc_dashboard-
End-to-end data platform integrating Google Sheets, Fieldwire, and accounting systems into Supabase and Laravel dashboards.
★ 0 6mo agoExplain → -
hands-on-aws-certified-data-engineer-associate-DEA-C01-Practice ⑂
No description.
★ 0 1y agoExplain → -
Data-Structures-and-Algorithms ⑂
Data Structures and Algorithms in Python
Python ★ 0 6mo agoExplain → -
data-engineering-zoomcamp ⑂
Data Engineering Zoomcamp is a free nine-week course that covers the fundamentals of data engineering.
★ 0 1y agoExplain → -
infinity_bh_appfolio
No description.
Python ★ 0 1y agoExplain → -
StreamlitDashboard
No description.
Python ★ 0 1y agoExplain → -
ur-found
The UR Found project was developed as part of a university initiative to improve item tracking and retrieval. It provides an efficient and user-friendly system for reporting, searching, and claiming lost items. This application features a centralized database, advanced search capabilities, and role-based access control for users.
CSS ★ 0 1y agoExplain → -
End-to-end-Sentiment-Data-Project ⑂
A real-time data streaming and sentiment analysis pipeline using Apache Spark on Databricks, structured into bronze, silver, and gold stages. It features MLflow for model tracking and demonstrates handling large-scale data streams with advanced data transformation techniques.
Jupyter Notebook ★ 0 1y agoExplain → -
sql_problems
No description.
★ 0 1y agoExplain → -
awesome-public-real-time-datasets ⑂
A list of publicly available datasets with real-time data maintained by the team at bytewax.io
★ 0 2y agoExplain → -
DE_bootcamp
No description.
Jupyter Notebook ★ 0 1y agoExplain → -
data-engineer-handbook ⑂
This is a repo with links to everything you'd ever want to learn about data engineering
Jupyter Notebook ★ 0 1y agoExplain → -
little-book-of-pipelines ⑂
This repository goes over how to handle massive variety in data engineering
★ 0 3y agoExplain → -
Motion_Lab_Data_Integration
Developed a robust database and analytics platform using Python and SQL to manage and analyze complex biomechanical data, including EEG, EMG, and kinematic measurements.
Jupyter Notebook ★ 0 1y agoExplain → -
python-cheat-sheet ⑂
No description.
★ 0 2y agoExplain → -
data-scientist-handbook ⑂
This is a repo with links to everything you'd ever want to learn about data science
★ 0 2y agoExplain → -
TwoWars
Employing the Reddit Praw API, extracted comments, formulated prompts for stance analysis, utilized Llama2 for both training and prediction purposes.
Jupyter Notebook ★ 0 2y agoExplain → -
EDA-Open-Asteroid-Dataset
EDA of the open asteroid dataset utilizes R and statistical methods
★ 0 2y agoExplain → -
Data-Engineering-test ⑂
The resources of the preparation course for Databricks Data Engineer Associate certification exam
★ 0 2y agoExplain → -
EDA-Movie
No description.
Jupyter Notebook ★ 0 2y agoExplain → -
laravel-adminlte ⑂
laravel 10 adminlte 3.2.0 blade template
★ 0 2y agoExplain → -
d2l-en ⑂
Interactive deep learning book with multi-framework code, math, and discussions. Adopted at 500 universities from 70 countries including Stanford, MIT, Harvard, and Cambridge.
★ 0 2y agoExplain → -
ubtz-computer-registration
No description.
★ 0 2y agoExplain → -
ubtz-data-exchange-info
No description.
Blade ★ 0 2y agoExplain → -
ubtz-person-on-railway
No description.
CSS ★ 0 2y agoExplain → -
ubtz-locomotive-repair
No description.
HTML ★ 0 2y agoExplain → -
ubtz-procurement
No description.
JavaScript ★ 0 2y agoExplain → -
ubtz-apartment
No description.
HTML ★ 0 2y agoExplain → -
ubtz-tender
No description.
Blade ★ 0 2y agoExplain → -
ubtz-locomotive-lent
No description.
Blade ★ 0 2y agoExplain → -
ubtz-wagon-management
No description.
PHP ★ 0 2y agoExplain → -
ubtz-autobaaz
No description.
Blade ★ 0 2y agoExplain → -
ubtz-container-management
No description.
Blade ★ 0 2y agoExplain → -
ubtz-data-exchange-count
No description.
PHP ★ 0 2y agoExplain → -
ubtz-locomotive-passport
No description.
HTML ★ 0 2y agoExplain → -
ubtz-wagon-sequence
No description.
PHP ★ 0 2y agoExplain → -
shop-management
No description.
JavaScript ★ 0 2y agoExplain → -
ubtz-locomotive-front
No description.
JavaScript ★ 0 2y agoExplain → -
ubtz-daily-info
No description.
Blade ★ 0 2y agoExplain → -
ubtz-xcdb
No description.
PHP ★ 0 2y agoExplain → -
ubtz-ztus-medeelel
No description.
HTML ★ 0 2y agoExplain →
No repos match these filters.