Category Β· 47 repos
Data Engineering & Analytics
ETL, warehouses, streaming, BI dashboards and data visualization. Ranked by star velocity over the last 24 hours.
Event Driven Orchestration & Scheduling Platform for Mission Critical Applications
Data Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. Join the course here ππΌ
πͺ Data Formulator is an interactive AI-powered data analysis system makes it easy to connect, explore and visualize data.
RustFS is an open-source, S3-compatible high-performance object storage system supporting migration and coexistence with other S3-compatible platforms such as MinIO and Ceph.
Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more
App to easily query, script, and visualize data from every database, file, and API.
π Welcome open-source Python mini-project contributions!
A lightweight opinionated ETL framework, halfway between plain scripts and Apache Airflow
Python web scrapers. Current collection: NetEase Cloud Music song scraper, Bilibili video scraper, Zhihu Q&A scraper, wallpaper scraper, xvideos video scraper, audiobook scraper, Weibo scraper, Anjuke information scraper + data visualization, Bilibili video-cover extractor, IP proxy-pool wrapper, large-scale Zhihu user scraper + data analysis, and GitHub user scraper.
From "Math Isn't Hard": the two-volume γLinear Algebra Isn't Hardγ, covering 66 topics; feedback is welcome.
Clinvis Data Visualization Unit For PTSD Explorer
Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does.
PixelShop dataset (DuckDB) for an introductory Business Intelligence course
A end-to-end MLOps pipeline for predicting telecom customer churn, featuring automated data preprocessing, ML model training, experiment tracking with MLflow, distributed training using PySpark, real-time inference via Kafka streaming, Airflow DAG orchestration, and Dockerized REST API deployment.
The Open Source Feature Store for AI/ML
ποΈ A comprehensive public archive of tweets by top hands-on traders on Twitter Β· Strip away the noise and distill a real-world trading knowledge system and fully open-source data lake, completely open source and free by ζ’ε.AI
Postgres intelligence for ai agents
This is a repo with links to everything you'd ever want to learn about data engineering
Implementing Event Driven Architecture and use Domain Driven Design approach. Technology: Go, Kafka, Docker.
A Claude Skill for creating academic presentations (conference talks, seminar slides, thesis defenses, grant briefings). Enforces action titles, structured argument, exhibit discipline, citation standards, and communication-first design. Works alongside Anthropic's built-in PPTX skill.
A big data knowledge repository covering data warehouse modeling, real-time computing, big data, data platforms, system design, Java, algorithms, and more.
All-in-one IoT Platform - Device management, data collection, processing and visualization.
Performant financial charts built with HTML5 canvas
Make your own running home page
WeFlow - a local app for exporting WeChat chat history and generating annual reports
Typed pipelines for Node.js. Define dependencies, preview execution, and inspect step results.
MetricFlow allows you to define, build, and maintain metrics in code.
Data visualization Skill for AI Agents, turning data into polished, interactive HTML charts. A data visualization Skill for AI Agents that quickly turns data into polished, interactive HTML charts.
10 Weeks, 20 Lessons, Data Science for All!
:hedgehog: PostHog is the leading platform for building self-driving products. Our developer tools β AI observability, analytics, session replay, flags, experiments, error tracking, logs, and more β capture all the context agents need to diagnose problems, uncover opportunities, and ship fixes. Stee
The missing visual surface companion for agents - generate UI mockups, data visualizations, code explainers, and more
Stream PostgreSQL changes into standard Apache Iceberg tables, built in Rust.
Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 250,000+ scientists worldwide. 177 ready-to-use validated skills plus 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Pi, Antigrav
Interactive Data Visualization in the browser, from Python
The HTML5 Creation Engine: Create beautiful digital content with the fastest, most flexible 2D WebGL renderer.
600+ free patterns and concepts for data engineers on Azure Databricks, Microsoft Fabric, and PySpark. 12 books covering the full stack. Free forever.
HR Attrition & Workforce Analysis using Excel and Tableau. Explores employee turnover patterns, demographics, and retention insights through interactive dashboards and pivot-based reporting.
Daniel15568/percentify β trending on GitHub.
OpenPanel is an open-source web and product analytics platform, an open-source alternative to Mixpanel with optional self-hosting.
Kafka broker embedded in postgres
:book: [Translation] Python for Data Analysis Β· 2nd Edition
The backtesting engine that gives you an unfair advantage. Run thousands of trading ideas before others finish one.
Notebooks for financial economics. Keywords: Jupyter notebook pandas Federal Reserve FRED Ferbus GDP CPI PCE inflation unemployment wage income debt Case-Shiller housing asset portfolio equities SPX bonds TIPS rates currency FX euro EUR USD JPY yen XAU gold Brent WTI oil Holt-Winters time-series for
Extension for Scikit-learn is a seamless way to speed up your Scikit-learn application
Alluxio, data orchestration for analytics and machine learning in the cloud
π Python training in French, from beginner to advanced | 13 modules: OOP, asyncio, testing, FastAPI, SQLAlchemy, Data Science (NumPy/Pandas) | Python 3.12+ | MIT
Aim π« β An easy-to-use & supercharged open-source experiment tracker.