Awesome Open Source
Awesome Open Source

= Awesome Open-Source Data Engineering :toc: :toc-placement!:

This[Awesome List] aims at providing an overview of[open-source] projects related to data engineering. This is a community effort: please[contribute] and send your pull requests for growing this list! For a list including non-OSS tools, see this amazing[Awesome List].


== Analytics

  •[Apache Spark] - A unified analytics engine for large-scale data processing. Includes APIs in Scala, Java, Python (known as PySpark), and R (SparkR).
  •[Apache Beam] - An open-source implementation of Google DataFlow. Provides capabilites of batch and streaming data processing jobs that run on any execution engine, including Spark, Flink, or its own DirectRunner. Supports multiple APIs in Java, Python, and Go.
  •[Apache Flink] - Stateful computations over data streams.
  •[Trino (formerly known as PrestoSQL)] - Distributed SQL Query Engine for Big Data.

== Business Intelligence

== Change Data Capture

== Datastores

== Data Governance and Registries

== Data Virtualization

== Data Orchestration

  •[Alluxio] - Scalable, multi-tiered distributed caching for HDFS, S3, Ceph, NFS, and related filestores. Provides integrations for SQL queries into a Catalog from Spark, Hive, and Presto.

== Formats

== Integration

== Messaging Infrastructure

== Specifications and Standards

== Stream Processing

== Testing

== Versioning

== Workflow Management

== Related Resources

only overview contents, no specific tools

=== Slide Decks, Recordings and Podcasts

=== Blog Posts and Articles

=== Collections

== License

The contents of this repository is licensed under the "Creative Commons Attribution-ShareAlike 4.0 International License".

Get A Weekly Email With Trending Projects For These Topics
No Spam. Unsubscribe easily at any time.
awesome-list (1,307
data-engineering (52