Apache Pig
A platform for analyzing large data sets with a high-level language and infrastructure for
Apache Pig is a platform for analyzing large data sets, built around Pig Latin, a high-level textual language for expressing data analysis programs. It addresses the need to write and maintain complex data transformations without handling the underlying distributed execution details. Pig programs are structured as data flow sequences, which makes them amenable to parallelization and allows the system to optimize their execution automatically.
Pig Latin provides a query algebra for operations such as merging data sets, filtering them, and applying functions to records or groups of records. Users can create their own functions for special-purpose processing. The infrastructure layer includes a compiler that produces sequences of Map-Reduce programs, executed on a Hadoop cluster. The current release supports Hadoop 3, Tez 0.10, Hive 3, Spark 3, HBase 2 and Python 3.
Apache Pig is an open source project released under the Apache 2.0 License, developed as a volunteer effort under the Apache Software Foundation. It is intended for users working with large-scale data analysis on distributed clusters.
12 alternatives to Apache Pig
Ranked by how well each tool replaces Apache Pig: shared features, audience, price and popularity.
Distributed data warehouse for massive-scale analytics.
Covers 4 of 6 key features.
Open source70 out of 100 match—- 62 out of 100 match—
Free, open repository of web crawl data usable by anyone.
Covers 3 of 6 key features and has a free plan.
Free plan60 out of 100 matchFreeOpen-source software for reliable, scalable, distributed computing
Covers 1 of 6 key features and has a free plan.
Free planOpen source59 out of 100 matchFreeA platform purpose-built for high-speed data engineering.
Covers 4 of 6 key features and has a free plan.
Free planOpen source58 out of 100 matchFree- 58 out of 100 matchUsage-based
- 57 out of 100 matchFree
Open source, native analytic database for open data and table formats.
Covers 1 of 6 key features.
Open source56 out of 100 match—A unified programming model for defining and executing batch and streaming data processing
Covers 1 of 6 key features.
Open source55 out of 100 match—Enterprise-grade unified stream and batch processing engine with event-time windowing.
Covers 3 of 6 key features.
Open source54 out of 100 match—Distributed real-time computation system for processing unbounded data streams
Covers 2 of 6 key features and has a free plan.
Free planOpen source54 out of 100 matchFree