lanyuanxiaoyao/hudi - hudi - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
Alexey Kudinkin	e5a2baeed0	[HUDI-3549] Removing dependency on "spark-avro" (#4955 ) Hudi will be taking on promise for it bundles to stay compatible with Spark minor versions (for ex 2.4, 3.1, 3.2): meaning that single build of Hudi (for ex "hudi-spark3.2-bundle") will be compatible with ALL patch versions in that minor branch (in that case 3.2.1, 3.2.0, etc) To achieve that we'll have to remove (and ban) "spark-avro" as a dependency, which on a few occasions was the root-cause of incompatibility b/w consecutive Spark patch versions (most recently 3.2.1 and 3.2.0, due to this PR). Instead of bundling "spark-avro" as dependency, we will be copying over some of the classes Hudi depends on and maintain them along the Hudi code-base to make sure we're able to provide for the aforementioned guarantee. To workaround arising compatibility issues we will be applying local patches to guarantee compatibility of Hudi bundles w/in the Spark minor version branches. Following Hudi modules to Spark minor branches is currently maintained: "hudi-spark3" -> 3.2.x "hudi-spark3.1.x" -> 3.1.x "hudi-spark2" -> 2.4.x Following classes hierarchies (borrowed from "spark-avro") are maintained w/in these Spark-specific modules to guarantee compatibility with respective minor version branches: AvroSerializer AvroDeserializer AvroUtils Each of these classes has been correspondingly copied from Spark 3.2.1 (for 3.2.x branch), 3.1.2 (for 3.1.x branch), 2.4.4 (for 2.4.x branch) into their respective modules. SchemaConverters class in turn is shared across all those modules given its relative stability (there're only cosmetical changes from 2.4.4 to 3.2.1). All of the aforementioned classes have their corresponding scope of visibility limited to corresponding packages (org.apache.spark.sql.avro, org.apache.spark.sql) to make sure broader code-base does not become dependent on them and instead relies on facades abstracting them. Additionally, given that Hudi plans on supporting all the patch versions of Spark w/in aforementioned minor versions branches of Spark, additional build steps were added to validate that Hudi could be properly compiled against those versions. Testing, however, is performed against the most recent patch versions of Spark with the help of Azure CI. Brief change log: - Removing spark-avro bundling from Hudi by default - Scaffolded Spark 3.2.x hierarchy - Bootstrapped Spark 3.1.x Avro serializer/deserializer hierarchy - Bootstrapped Spark 2.4.x Avro serializer/deserializer hierarchy - Moved ExpressionCodeGen,ExpressionPayload into hudi-spark module - Fixed AvroDeserializer to stay compatible w/ both Spark 3.2.1 and 3.2.0 - Modified bot.yml to build full matrix of support Spark versions - Removed "spark-avro" dependency from all modules - Fixed relocation of spark-avro classes in bundles to assist in running integ-tests.	2022-03-29 14:44:47 -04:00
leesf	8f8a8158e2	[HUDI-2520] Fix drop table issue when sync to Hive (#5143 )	2022-03-28 19:34:12 -07:00
Alexey Kudinkin	8b38ddedc2	[HUDI-3594] Supporting Composite Expressions over Data Table Columns in Data Skipping flow (#4996 )	2022-03-24 22:27:15 -07:00
ForwardXu	08fd80c913	[HUDI-3221] Support querying a table as of a savepoint (#4720 )	2022-03-08 10:02:34 -08:00
Alexey Kudinkin	85e8a5c4de	[HUDI-1296] Support Metadata Table in Spark Datasource (#4789 ) * Bootstrapping initial support for Metadata Table in Spark Datasource - Consolidated Avro/Row conversion utilities to center around Spark's AvroDeserializer ; removed duplication - Bootstrapped HoodieBaseRelation - Updated HoodieMergeOnReadRDD to be able to handle Metadata Table - Modified MOR relations to be able to read different Base File formats (Parquet, HFile)	2022-02-24 16:23:13 -05:00
leesf	2a93b8efb2	[HUDI-3489] Unify config to avoid duplicate code (#4883 )	2022-02-23 08:14:30 -05:00
leesf	0db1e978c6	[HUDI-3254] Introduce HoodieCatalog to manage tables for Spark Datasource V2 (#4611 )	2022-02-14 06:26:58 -08:00
leesf	5ce45c440b	[HUDI-3172] Refactor hudi existing modules to make more code reuse in V2 Implementation (#4514 ) * Introduce hudi-spark3-common and hudi-spark2-common modules to place classes that would be reused in different spark versions, also introduce hudi-spark3.1.x to support spark 3.1.x. * Introduce hudi format under hudi-spark2, hudi-spark3, hudi-spark3.1.x modules and change the hudi format in original hudi-spark module to hudi_v1 format. * Manually tested on Spark 3.1.2 and Spark 3.2.0 SQL. * Added a README.md file under hudi-spark-datasource module.	2022-01-14 13:42:35 +08:00
leesf	29ab6fb9ad	[HUDI-3140] Fix bulk_insert failure on Spark 3.2.0 (#4498 )	2022-01-04 09:59:59 +08:00
Yann Byron	05942e018c	[HUDI-2811] Support Spark 3.2 (#4270 )	2021-12-28 00:12:44 -08:00
xiarixiaoyao	9246b16492	[HUDI-2958] Automatically set spark.sql.parquet.writelegacyformat, when using bulkinsert to insert data which contains decimalType (#4253 )	2021-12-17 08:58:02 -05:00
yuzhao.cyz	a1d0ff4209	Moving to 0.11.0-SNAPSHOT on master branch.	2021-11-27 17:22:10 +08:00
Raymond Xu	02f7ca2b05	[HUDI-1870] Add more Spark CI build tasks (#4022 ) * [HUDI-1870] Add more Spark CI build tasks - build for spark3.0.x - build for spark-shade-unbundle-avro - fix build failures - delete unnecessary assertion for spark 3.0.x - use AvroConversionUtils#convertAvroSchemaToStructType instead of calling SchemaConverters#toSqlType directly to solve the compilation failures with spark-shade-unbundle-avro (#5) Co-authored-by: Yann <biyan900116@gmail.com>	2021-11-22 02:16:45 -08:00
Yann Byron	1f17467f73	[HUDI-1869] Upgrading Spark3 To 3.1 (#3844 ) Co-authored-by: pengzhiwei <pengzhiwei2015@icloud.com>	2021-11-02 18:25:12 -07:00
董可伦	2f07e1267f	[MINOR] Fix typo Hooodie corrected to Hoodie & reuqired corrected to required (#3730 )	2021-09-30 09:55:32 +08:00
Udit Mehrotra	c350d05dd3	Restore 0.8.0 config keys with deprecated annotation (#3506 ) Co-authored-by: Sagar Sumit <sagarsumit09@gmail.com> Co-authored-by: Vinoth Chandar <vinoth@apache.org>	2021-08-19 13:36:40 -07:00
Udit Mehrotra	3e301196bf	Moving to 0.10.0-SNAPSHOT on master branch.	2021-08-14 18:51:09 -07:00
zhangyue19921010	73d898322b	[MINOR] Fix travis from errors (#3432 )	2021-08-10 08:25:49 -07:00
wenningd	91bb0d1318	[HUDI-2255] Refactor Datasource options (#3373 ) Co-authored-by: Wenning Ding <wenningd@amazon.com>	2021-08-03 17:50:30 -07:00
Sivabalan Narayanan	7bdae69053	[HUDI-2253] Refactoring few tests to reduce runningtime. DeltaStreamer and MultiDeltaStreamer tests. Bulk insert row writer tests (#3371 ) Co-authored-by: Sivabalan Narayanan <nsb@Sivabalans-MBP.attlocal.net>	2021-07-29 22:22:26 -07:00
Vinay Patil	5a94b6bf54	[HUDI-2192] Clean up Multiple versions of scala libraries detected Warning (#3292 )	2021-07-21 00:33:27 -07:00
Sivabalan Narayanan	d5026e9a24	[HUDI-2161] Adding support to disable meta columns with bulk insert operation (#3247 )	2021-07-19 20:43:48 -04:00
pengzhiwei	572a214412	[HUDI-1884] MergeInto Support Partial Update For COW (#3154 )	2021-07-17 12:59:18 +08:00
vinoth chandar	75040ee9e5	[HUDI-2149] Ensure and Audit docs for every configuration class in the codebase (#3272 ) - Added docs when missing - Rewrote, reworded as needed - Made couple more classes extend HoodieConfig	2021-07-14 10:56:08 -07:00
Sivabalan Narayanan	8c0dbaa9b3	[HUDI-2009] Fixing extra commit metadata in row writer path (#3075 )	2021-07-08 03:07:27 -04:00
Sivabalan Narayanan	ea9e5d0e8b	[HUDI-1104] Adding support for UserDefinedPartitioners and SortModes to BulkInsert with Rows (#3149 )	2021-07-07 11:15:25 -04:00
wenningd	d412fb2fe6	[HUDI-89] Add configOption & refactor all configs based on that (#2833 ) Co-authored-by: Wenning Ding <wenningd@amazon.com>	2021-06-30 14:26:30 -07:00
pengzhiwei	f760ec543e	[HUDI-1659] Basic Implement Of Spark Sql Support For Hoodie (#2645 ) Main functions: Support create table for hoodie. Support CTAS. Support Insert for hoodie. Including dynamic partition and static partition insert. Support MergeInto for hoodie. Support DELETE Support UPDATE Both support spark2 & spark3 based on DataSourceV1. Main changes: Add sql parser for spark2. Add HoodieAnalysis for sql resolve and logical plan rewrite. Add commands implementation for CREATE TABLE、INSERT、MERGE INTO & CTAS. In order to push down the update&insert logical to the HoodieRecordPayload for MergeInto, I make same change to the HoodieWriteHandler and other related classes. 1、Add the inputSchema for parser the incoming record. This is because the inputSchema for MergeInto is different from writeSchema as there are some transforms in the update& insert expression. 2、Add WRITE_SCHEMA to HoodieWriteConfig to pass the write schema for merge into. 3、Pass properties to HoodieRecordPayload#getInsertValue to pass the insert expression and table schema. Verify this pull request Add TestCreateTable for test create hoodie tables and CTAS. Add TestInsertTable for test insert hoodie tables. Add TestMergeIntoTable for test merge hoodie tables. Add TestUpdateTable for test update hoodie tables. Add TestDeleteTable for test delete hoodie tables. Add TestSqlStatement for test supported ddl/dml currently.	2021-06-07 23:24:32 -07:00
wangxianghu	f3777f44fe	[MINOR] Remove unused imports and some other checkstyle issues (#2800 )	2021-04-11 21:42:34 +08:00
pengzhiwei	684622c7c9	[HUDI-1591] Implement Spark's FileIndex for Hudi to support queries via Hudi DataSource using non-globbed table path and partition pruning (#2651 )	2021-04-01 11:12:28 -07:00
Gary Li	452f5e2d66	[HOTFIX] close spark session in functional test suite and disable spark3 test for spark2 (#2727 )	2021-03-29 06:04:48 -07:00
garyli1019	6e803e08b1	Moving to 0.9.0-SNAPSHOT on master branch.	2021-03-24 21:37:14 +08:00
Sivabalan Narayanan	b038623ed3	[HUDI 1615] Fixing null schema in bulk_insert row writer path (#2653 ) * [HUDI-1615] Avoid passing in null schema from row writing/deltastreamer * Fixing null schema in bulk insert row writer path * Fixing tests Co-authored-by: vc <vinoth@apache.org>	2021-03-16 09:44:11 -07:00
pengzhiwei	37972071ff	[HUDI-1109] Support Spark Structured Streaming read from Hudi table (#2485 )	2021-02-17 03:36:29 -08:00
Vinoth Chandar	3719e7b388	Moving to 0.8.0-SNAPSHOT on master branch.	2021-01-20 11:31:22 -08:00
Sivabalan Narayanan	b9c2856d16	[HUDI-1535] Fix 0.7.0 snapshot (#2456 ) * Revert "[MINOR] Bumping snapshot version to 0.7.0 (#2435)" This reverts commit `a43e191d6c`. * Fixing 0.7.0 snapshot bump	2021-01-19 12:20:43 -08:00
Sivabalan Narayanan	a43e191d6c	[MINOR] Bumping snapshot version to 0.7.0 (#2435 )	2021-01-16 09:56:28 -05:00
wangxianghu	b593f10629	[MINOR] Rename unit test package of hudi-spark3 from scala to java (#2411 )	2021-01-06 23:07:24 +08:00
wenningd	286055ce34	[HUDI-1451] Support bulk insert v2 with Spark 3.0.0 (#2328 ) Co-authored-by: Wenning Ding <wenningd@amazon.com> - Added support for bulk insert v2 with datasource v2 api in Spark 3.0.0.	2020-12-25 09:43:34 -05:00
wenningd	fce1453fa6	[HUDI-1040] Make Hudi support Spark 3 (#2208 ) * Fix flaky MOR unit test * Update Spark APIs to make it be compatible with both spark2 & spark3 * Refactor bulk insert v2 part to make Hudi be able to compile with Spark3 * Add spark3 profile to handle fasterxml & spark version * Create hudi-spark-common module & refactor hudi-spark related modules Co-authored-by: Wenning Ding <wenningd@amazon.com>	2020-12-09 15:52:23 -08:00

40 Commits