Go to file

pengzhiwei f760ec543e [HUDI-1659] Basic Implement Of Spark Sql Support For Hoodie (#2645 )

Main functions:
Support create table for hoodie.
Support CTAS.
Support Insert for hoodie. Including dynamic partition and static partition insert.
Support MergeInto for hoodie.
Support DELETE
Support UPDATE
Both support spark2 & spark3 based on DataSourceV1.

Main changes:
Add sql parser for spark2.
Add HoodieAnalysis for sql resolve and logical plan rewrite.
Add commands implementation for CREATE TABLE、INSERT、MERGE INTO & CTAS.
In order to push down the update&insert logical to the HoodieRecordPayload for MergeInto, I make same change to the
HoodieWriteHandler and other related classes.
1、Add the inputSchema for parser the incoming record. This is because the inputSchema for MergeInto is different from writeSchema as there are some transforms in the update& insert expression.
2、Add WRITE_SCHEMA to HoodieWriteConfig to pass the write schema for merge into.
3、Pass properties to HoodieRecordPayload#getInsertValue to pass the insert expression and table schema.


Verify this pull request
Add TestCreateTable for test create hoodie tables and CTAS.
Add TestInsertTable for test insert hoodie tables.
Add TestMergeIntoTable for test merge hoodie tables.
Add TestUpdateTable for test update hoodie tables.
Add TestDeleteTable for test delete hoodie tables.
Add TestSqlStatement for test supported ddl/dml currently.

2021-06-07 23:24:32 -07:00

.github

[HUDI-985] Introduce rerun ci bot (#1693 )

2020-07-22 22:59:24 -07:00

docker

[HUDI-1851] Adding test suite long running automate scripts for docker (#2880 )

2021-05-11 01:26:01 -07:00

hudi-cli

[HUDI-1914] Add fetching latest schema to table command in hudi-cli (#2964 )

2021-06-07 16:04:35 -07:00

hudi-client

[HUDI-1659] Basic Implement Of Spark Sql Support For Hoodie (#2645 )

2021-06-07 23:24:32 -07:00

hudi-common

[HUDI-1659] Basic Implement Of Spark Sql Support For Hoodie (#2645 )

2021-06-07 23:24:32 -07:00

hudi-examples

[HUDI-1735] Add hive-exec dependency for hudi-examples (#2737 )

2021-03-30 21:35:16 +08:00

hudi-flink

add BootstrapFunction to support index bootstrap (#3024 )

2021-06-08 13:55:25 +08:00

hudi-hadoop-mr

[HUDI-1967] Fix the NPE for MOR Hive rt table query (#3032 )

2021-06-05 01:06:34 -07:00

hudi-integ-test

[HUDI-1659] Basic Implement Of Spark Sql Support For Hoodie (#2645 )

2021-06-07 23:24:32 -07:00

hudi-spark-datasource

[HUDI-1659] Basic Implement Of Spark Sql Support For Hoodie (#2645 )

2021-06-07 23:24:32 -07:00

hudi-sync

[HUDI-1659] Basic Implement Of Spark Sql Support For Hoodie (#2645 )

2021-06-07 23:24:32 -07:00

hudi-timeline-service

[HUDI-1707] Reduces log level for too verbose messages from info to debug level. (#2714 )

2021-05-10 07:16:02 -07:00

hudi-utilities

[HUDI-1659] Basic Implement Of Spark Sql Support For Hoodie (#2645 )

2021-06-07 23:24:32 -07:00

packaging

[HUDI-1659] Basic Implement Of Spark Sql Support For Hoodie (#2645 )

2021-06-07 23:24:32 -07:00

scripts

[MINOR] Add Missing Apache License to test files (#2736 )

2021-03-29 07:17:23 -07:00

style

[HUDI-1659] Basic Implement Of Spark Sql Support For Hoodie (#2645 )

2021-06-07 23:24:32 -07:00

.asf.yaml

[MINOR] Add apacheflink label (#2268 )

2020-11-22 10:41:11 +08:00

.codecov.yml

[HUDI-896] Report test coverage by modules & parallelize CI (#1753 )

2020-06-27 23:16:12 -07:00

.gitignore

[HUDI-985] Introduce rerun ci bot (#1693 )

2020-07-22 22:59:24 -07:00

.travis.yml

[MINOR] Make a separate travis CI job for hudi-utilities (#2469 )

2021-01-20 21:46:05 -08:00

azure-pipelines.yml

[HUDI-1620] Fix Metrics UT (#2894 )

2021-04-30 11:20:41 -07:00

doap_HUDI.rdf

[MINOR] Update doap with 0.8.0 release (#2772 )

2021-04-08 11:06:13 -04:00

LICENSE

[HUDI-1040] Make Hudi support Spark 3 (#2208 )

2020-12-09 15:52:23 -08:00

NOTICE

[HUDI-938] Removing incubating/incubator from project (#1658 )

2020-05-24 18:28:13 +08:00

pom.xml

[HUDI-1659] Basic Implement Of Spark Sql Support For Hoodie (#2645 )

2021-06-07 23:24:32 -07:00

README.md

[MINOR] Fix deprecated build link for travis (#2778 )

2021-04-07 08:57:10 +08:00

README.md

Apache Hudi

Apache Hudi (pronounced Hoodie) stands for Hadoop Upserts Deletes and Incrementals. Hudi manages the storage of large analytical datasets on DFS (Cloud stores, HDFS or any Hadoop FileSystem compatible storage).

https://hudi.apache.org/

Features

Upsert support with fast, pluggable indexing
Atomically publish data with rollback support
Snapshot isolation between writer & queries
Savepoints for data recovery
Manages file sizes, layout using statistics
Async compaction of row & columnar data
Timeline metadata to track lineage
Optimize data lake layout with clustering

Hudi supports three types of queries:

Snapshot Query - Provides snapshot queries on real-time data, using a combination of columnar & row-based storage (e.g Parquet + Avro).
Incremental Query - Provides a change stream with records inserted or updated after a point in time.
Read Optimized Query - Provides excellent snapshot query performance via purely columnar storage (e.g. Parquet).

Learn more about Hudi at https://hudi.apache.org

Building Apache Hudi from source

Prerequisites for building Apache Hudi:

Unix-like system (like Linux, Mac OS X)
Java 8 (Java 9 or 10 may work)
Git
Maven

# Checkout code and build
git clone https://github.com/apache/hudi.git && cd hudi
mvn clean package -DskipTests

# Start command
spark-2.4.4-bin-hadoop2.7/bin/spark-shell \
  --jars `ls packaging/hudi-spark-bundle/target/hudi-spark-bundle_2.11-*.*.*-SNAPSHOT.jar` \
  --conf 'spark.serializer=org.apache.spark.serializer.KryoSerializer'

To build the Javadoc for all Java and Scala classes:

# Javadoc generated under target/site/apidocs
mvn clean javadoc:aggregate -Pjavadocs

Build with Scala 2.12

The default Scala version supported is 2.11. To build for Scala 2.12 version, build using scala-2.12 profile

mvn clean package -DskipTests -Dscala-2.12

Build with Spark 3.0.0

The default Spark version supported is 2.4.4. To build for Spark 3.0.0 version, build using spark3 profile

mvn clean package -DskipTests -Dspark3

Build without spark-avro module

The default hudi-jar bundles spark-avro module. To build without spark-avro module, build using spark-shade-unbundle-avro profile

# Checkout code and build
git clone https://github.com/apache/hudi.git && cd hudi
mvn clean package -DskipTests -Pspark-shade-unbundle-avro

# Start command
spark-2.4.4-bin-hadoop2.7/bin/spark-shell \
  --packages org.apache.spark:spark-avro_2.11:2.4.4 \
  --jars `ls packaging/hudi-spark-bundle/target/hudi-spark-bundle_2.11-*.*.*-SNAPSHOT.jar` \
  --conf 'spark.serializer=org.apache.spark.serializer.KryoSerializer'

Running Tests

Unit tests can be run with maven profile unit-tests.

mvn -Punit-tests test

Functional tests, which are tagged with @Tag("functional"), can be run with maven profile functional-tests.

mvn -Pfunctional-tests test

To run tests with spark event logging enabled, define the Spark event log directory. This allows visualizing test DAG and stages using Spark History Server UI.

mvn -Punit-tests test -DSPARK_EVLOG_DIR=/path/for/spark/event/log

Quickstart

Please visit https://hudi.apache.org/docs/quick-start-guide.html to quickly explore Hudi's capabilities using spark-shell.

Languages

Java 81.4%

Scala 16.7%

ANTLR 0.9%

Shell 0.8%

Dockerfile 0.2%