lanyuanxiaoyao/hudi - hudi - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
pengzhiwei	0d8a4d0a56	[HUDI-1550] Honor ordering field for MOR Spark datasource reader (#2497 )	2021-02-01 21:04:27 +08:00
jiangjiguang	5d053b495b	[MINOR] Quickstart.generateUpdates method add check (#2505 )	2021-01-30 10:28:00 +08:00
satishkotha	2d2d5c83b1	[HUDI-1555] Remove isEmpty to improve clustering execution performance (#2502 )	2021-01-29 10:27:09 -08:00
wenningd	976420c49a	[HUDI-1512] Fix spark 2 unit tests failure with Spark 3 (#2412 ) * [HUDI-1512] Fix spark 2 unit tests failure with Spark 3 * resolve comments Co-authored-by: Wenning Ding <wenningd@amazon.com>	2021-01-21 07:04:28 -08:00
Vinoth Chandar	3719e7b388	Moving to 0.8.0-SNAPSHOT on master branch.	2021-01-20 11:31:22 -08:00
liujinhui	244f6def9c	[MINOR] Fix dataSource cannot use hoodie.datasource.hive_sync.auto_create_database (#2444 ) fix dataSource cannot use hoodie.datasource.hive_sync.auto_create_database	2021-01-20 22:58:18 +08:00
vinoth chandar	5ca0625b27	[HUDI 1308] Harden RFC-15 Implementation based on production testing (#2441 ) Addresses leaks, perf degradation observed during testing. These were regressions from the original rfc-15 PoC implementation. * Pass a single instance of HoodieTableMetadata everywhere * Fix tests and add config for enabling metrics - Removed special casing of assumeDatePartitioning inside FSUtils#getAllPartitionPaths() - Consequently, IOException is never thrown and many files had to be adjusted - More diligent handling of open file handles in metadata table - Added config for controlling reuse of connections - Added config for turning off fallback to listing, so we can see tests fail - Changed all ipf listing code to cache/amortize the open/close for better performance - Timelineserver also reuses connections, for better performance - Without timelineserver, when metadata table is opened from executors, reuse is not allowed - HoodieMetadataConfig passed into HoodieTableMetadata#create as argument. - Fix TestHoodieBackedTableMetadata#testSync	2021-01-19 21:20:28 -08:00
Sivabalan Narayanan	b9c2856d16	[HUDI-1535] Fix 0.7.0 snapshot (#2456 ) * Revert "[MINOR] Bumping snapshot version to 0.7.0 (#2435)" This reverts commit `a43e191d6c`. * Fixing 0.7.0 snapshot bump	2021-01-19 12:20:43 -08:00
Sivabalan Narayanan	a43e191d6c	[MINOR] Bumping snapshot version to 0.7.0 (#2435 )	2021-01-16 09:56:28 -05:00
lw0090	de42adc230	[HUDI-1520] add configure for spark sql overwrite use INSERT_OVERWRITE_TABLE (#2428 )	2021-01-11 09:07:47 -08:00
Udit Mehrotra	7ce3ac778e	[HUDI-1479] Use HoodieEngineContext to parallelize fetching of partiton paths (#2417 ) * [HUDI-1479] Use HoodieEngineContext to parallelize fetching of partition paths * Adding testClass for FileSystemBackedTableMetadata Co-authored-by: Nishith Agarwal <nagarwal@uber.com>	2021-01-10 21:19:52 -08:00
Gary Li	23e93d05c0	[MINOR] fix spark 3 build for incremental query on MOR (#2425 )	2021-01-09 21:08:55 -08:00
lw0090	368c1a8f5c	[HUDI-1399] support a independent clustering spark job to asynchronously clustering (#2379 ) * [HUDI-1481] add structured streaming and delta streamer clustering unit test * [HUDI-1399] support a independent clustering spark job to asynchronously clustering * [HUDI-1399] support a independent clustering spark job to asynchronously clustering * [HUDI-1498] Read clustering plan from requested file for inflight instant (#2389) * [HUDI-1399] support a independent clustering spark job with schedule generate instant time Co-authored-by: satishkotha <satishkotha@uber.com>	2021-01-09 17:30:16 -08:00
Gary Li	79ec7b4894	[HUDI-920] Support Incremental query for MOR table (#1938 )	2021-01-09 08:02:08 -08:00
Udit Mehrotra	17df517b81	[HUDI-1510] Move HoodieEngineContext and its dependencies to hudi-common (#2410 )	2021-01-07 11:34:06 -08:00
wangxianghu	b593f10629	[MINOR] Rename unit test package of hudi-spark3 from scala to java (#2411 )	2021-01-06 23:07:24 +08:00
Ryan Pifer	4b94529aaf	[HUDI-1325] [RFC-15] Merge updates of unsynced instants to metadata table (apache#2342) [RFC-15] Fix partition key in metadata table when bootstrapping from file system (apache#2387) Co-authored-by: Ryan Pifer <ryanpife@amazon.com>	2021-01-04 07:59:47 -08:00
Udit Mehrotra	4e64226844	[HUDI-1450] Use metadata table for listing in HoodieROTablePathFilter (apache#2326) [HUDI-1394] [RFC-15] Use metadata table (if present) to get all partition paths (apache#2351)	2021-01-04 07:59:47 -08:00
Gary Li	c5e8a024f6	[HUDI-1418] Set up flink client unit test infra (#2281 )	2020-12-31 08:57:22 +08:00
pengzhiwei	b83d1d3e61	[HUDI-1484] Escape the partition value in HiveSyncTool (#2363 )	2020-12-28 23:02:36 -05:00
lw0090	9e6889a8ce	[HUDI-1481] add structured streaming and delta streamer clustering unit test (#2360 )	2020-12-27 20:27:09 -08:00
lw0090	e807bb895e	[HUDI-1487] fix unit test testCopyOnWriteStorage random failed (#2364 )	2020-12-25 09:54:23 -08:00
wenningd	286055ce34	[HUDI-1451] Support bulk insert v2 with Spark 3.0.0 (#2328 ) Co-authored-by: Wenning Ding <wenningd@amazon.com> - Added support for bulk insert v2 with datasource v2 api in Spark 3.0.0.	2020-12-25 09:43:34 -05:00
wenningd	89f482eaf2	[HUDI-1489] Fix null pointer exception when reading updated written bootstrap table (#2370 ) Co-authored-by: Wenning Ding <wenningd@amazon.com>	2020-12-23 11:26:24 -08:00
wangxianghu	f8ccb2872d	[HUDI-1471] Make QuickStartUtils generate deletes according to specific ts (#2357 )	2020-12-22 21:14:18 +08:00
Sivabalan Narayanan	33d338f392	[HUDI-115] Adding DefaultHoodieRecordPayload to honor ordering with combineAndGetUpdateValue (#2311 ) * Added ability to pass in `properties` to payload methods, so they can perform table/record specific merges * Added default methods so existing payload classes are backwards compatible. * Adding DefaultHoodiePayload to honor ordering while merging two records * Fixing default payload based on feedback	2020-12-19 19:19:42 -08:00
lw0090	8b5d6f9430	[HUDI-1437] support more accurate spark JobGroup for better performance tracking (#2322 )	2020-12-17 15:20:13 -08:00
wangxianghu	4ddfc61d70	[MINOR] Make QuickstartUtil generate random timestamp instead of 0 (#2340 )	2020-12-17 18:00:23 +08:00
wenningd	26cdc457f6	[HUDI-1376] Drop Hudi metadata cols at the beginning of Spark datasource writing (#2233 ) Co-authored-by: Wenning Ding <wenningd@amazon.com>	2020-12-15 16:20:48 -08:00
wangxianghu	6cf25d5c8a	[MINOR] Minor improve in IncrementalRelation (#2314 )	2020-12-10 20:16:00 +08:00
Danny Chan	4bc45a391a	[HUDI-1445] Refactor AbstractHoodieLogRecordScanner to use Builder (#2313 )	2020-12-10 20:02:02 +08:00
wenningd	fce1453fa6	[HUDI-1040] Make Hudi support Spark 3 (#2208 ) * Fix flaky MOR unit test * Update Spark APIs to make it be compatible with both spark2 & spark3 * Refactor bulk insert v2 part to make Hudi be able to compile with Spark3 * Add spark3 profile to handle fasterxml & spark version * Create hudi-spark-common module & refactor hudi-spark related modules Co-authored-by: Wenning Ding <wenningd@amazon.com>	2020-12-09 15:52:23 -08:00

32 Commits