lanyuanxiaoyao/hudi - hudi - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
Roc Marshal	fcedbfcb58	[MINOR][hudi-client] Code-cleanup,remove redundant variable declarations (#2956 )	2021-05-17 13:34:42 +08:00
Danny Chan	8869b3b418	[HUDI-1902] Clean the corrupted files generated by FlinkMergeAndReplaceHandle (#2949 ) Make the intermediate files of FlinkMergeAndReplaceHandle hidden, when committing the instant, clean these files in case there was some corrupted files left(in normal case, the intermediate files should be cleaned by the FlinkMergeAndReplaceHandle itself).	2021-05-14 15:43:37 +08:00
xoln ann	12443e4187	[HUDI-1446] Support skip bootstrapIndex's init in abstract fs view init (#2520 ) Co-authored-by: zhongliang <zhongliang@kuaishou.com> Co-authored-by: Sivabalan Narayanan <sivabala@uber.com>	2021-05-14 00:29:26 -04:00
Danny Chan	ad77cf42ba	[HUDI-1900] Always close the file handle for a flink mini-batch write (#2943 ) Close the file handle eagerly to avoid corrupted files as much as possible.	2021-05-14 10:25:18 +08:00
Danny Chan	b98c9ab439	[HUDI-1895] Close the file handles gracefully for flink write function to avoid corrupted files (#2938 )	2021-05-12 18:44:10 +08:00
lw0090	5a8b2a4f86	[HUDI-1768] add spark datasource unit test for schema validate add column (#2776 )	2021-05-11 16:49:18 -04:00
TeRS-K	be9db2c4f5	[HUDI-1055] Remove hardcoded parquet in tests (#2740 ) * Remove hardcoded parquet in tests * Use DataFileUtils.getInstance * Renaming DataFileUtils to BaseFileUtils Co-authored-by: Vinoth Chandar <vinoth@apache.org>	2021-05-11 10:01:45 -07:00
Danny Chan	42ec7e30d7	[HUDI-1890] FlinkCreateHandle and FlinkAppendHandle canWrite should always return true (#2933 ) The method #canWrite should always return true because they can already write based on file size, e.g. the BucketAssigner.	2021-05-11 09:14:51 +08:00
Danny Chan	c1b331bcff	[HUDI-1886] Avoid to generates corrupted files for flink sink (#2929 )	2021-05-10 10:43:03 +08:00
Danny Chan	bfbf993cbe	[HUDI-1878] Add max memory option for flink writer task (#2920 ) Also removes the rate limiter because it has the similar functionality, modify the create and merge handle cleans the retry files automatically.	2021-05-08 14:27:56 +08:00
Danny Chan	528f4ca988	[HUDI-1880] Support streaming read with compaction and cleaning (#2921 )	2021-05-07 20:04:35 +08:00
Sivabalan Narayanan	0284cdecce	[HUDI-1876] wiring in Hadoop Conf with AvroSchemaConverters instantiation (#2914 )	2021-05-05 21:31:44 -07:00
Raymond Xu	3418a92de8	[HUDI-1620] Fix Metrics UT (#2894 ) Make sure shutdown Metrics between unit test cases to ensure isolation	2021-04-30 11:20:41 -07:00
satishkotha	386767693d	[HUDI-1833] rollback pending clustering even if there is greater commit (#2863 ) * [HUDI-1833] rollback pending clustering even if there are greater commits	2021-04-27 14:21:42 -07:00
satishkotha	2999586509	[HUDI-1690] use jsc union instead of rdd union (#2872 )	2021-04-26 23:35:01 -07:00
Roc Marshal	9bbb458e88	[MINOR] Remove redundant method-calling. (#2881 )	2021-04-27 09:34:09 +08:00
Danny Chan	d047e91d86	[HUDI-1837] Add optional instant range to log record scanner for log (#2870 )	2021-04-26 16:53:18 +08:00
Chanh Le	a1e636dc6b	[HUDI-1551] Add support for BigDecimal and Integer when partitioning based on time. (#2851 ) Co-authored-by: trungchanh.le <trungchanh.le@bybit.com>	2021-04-22 21:56:20 +08:00
jsbali	b31c520c66	[HUDI-1714] Added tests to TestHoodieTimelineArchiveLog for the archival of compl… (#2677 ) * Added tests to TestHoodieTimelineArchiveLog for the archival of completed clean and rollback actions. * Adding code review changes * [HUDI-1714] Minor Fixes	2021-04-21 10:27:43 -07:00
Sebastian Bernauer	9a288ccbeb	[MINOR] Added metric reporter Prometheus to HoodieBackedTableMetadataWriter (#2842 )	2021-04-19 16:04:59 -07:00
li36909	6b4b878d08	[HUDI-1744] rollback fails on mor table when the partition path hasn't any files (#2749 ) Co-authored-by: lrz <lrz@lrzdeMacBook-Pro.local>	2021-04-19 15:44:11 -07:00
Aditya Tiwari	ec2334ceac	[HUDI-1716]: Resolving default values for schema from dataframe (#2765 ) - Adding default values and setting null as first entry in UNION data types in avro schema. Co-authored-by: Aditya Tiwari <aditya.tiwari@flipkart.com>	2021-04-19 10:05:20 -04:00
Danny Chan	dab5114f16	[HUDI-1804] Continue to write when Flink write task restart because of container killing (#2843 ) The `FlinkMergeHande` creates a marker file under the metadata path each time it initializes, when a write task restarts from killing, it tries to create the existing file and reports error. To solve this problem, skip the creation and use the original data file as base file to merge.	2021-04-19 19:43:41 +08:00
Danny Chan	b6d949b48a	[HUDI-1801] FlinkMergeHandle rolling over may miss to rename the latest file handle (#2831 ) The FlinkMergeHandle may rename the N-1 th file handle instead of the latest one, thus to cause data duplication.	2021-04-16 11:40:53 +08:00
Danny Chan	ab4a7b0b4a	[HUDI-1788] Insert overwrite (table) for Flink writer (#2808 ) Supports `INSERT OVERWRITE` and `INSERT OVERWRITE TABLE` for Flink writer.	2021-04-14 10:23:37 +08:00
wangxianghu	040756d8c0	[HUDI-1785] Move OperationConverter to hudi-client-common for code reuse (#2798 )	2021-04-12 16:22:33 +08:00
hj2016	1da16dfd2e	[HUDI-1784] Added print detailed stack log when hbase connection error (#2799 )	2021-04-12 13:46:06 +08:00
hongdd	ecdbd2517f	[HUDI-699] Fix CompactionCommand and add unit test for CompactionCommand (#2325 )	2021-04-08 15:35:33 +08:00
Simon	18459d4045	[MINOR] Some unit test code optimize (#2782 ) * Optimized code * Optimized code	2021-04-08 13:35:03 +08:00
Danny Chan	9c369c607d	[HUDI-1757] Assigns the buckets by record key for Flink writer (#2757 ) Currently we assign the buckets by record partition path which could cause hotspot if the partition field is datetime type. Changes to assign buckets by grouping the record whth their key first, the assignment is valid if only there is no conflict(two task write to the same bucket). This patch also changes the coordinator execution to be asynchronous.	2021-04-06 19:06:41 +08:00
Roc Marshal	94a5e72f16	[HUDI-1737][hudi-client] Code Cleanup: Extract common method in HoodieCreateHandle & FlinkCreateHandle (#2745 )	2021-04-02 11:39:05 +08:00
pengzhiwei	684622c7c9	[HUDI-1591] Implement Spark's FileIndex for Hudi to support queries via Hudi DataSource using non-globbed table path and partition pruning (#2651 )	2021-04-01 11:12:28 -07:00
Danny Chan	9804662bc8	[HUDI-1738] Emit deletes for flink MOR table streaming read (#2742 ) Current we did a soft delete for DELETE row data when writes into hoodie table. For streaming read of MOR table, the Flink reader detects the delete records and still emit them if the record key semantics are still kept. This is useful and actually a must for streaming ETL pipeline incremental computation.	2021-04-01 15:25:31 +08:00
vinoyang	fe16d0de7c	[MINOR] Delete useless UpsertPartitioner for flink integration (#2746 )	2021-03-31 16:36:42 +08:00
Sebastian Bernauer	aa0da72c59	Preparation for Avro update (#2650 )	2021-03-30 21:50:17 -07:00
leo-Iamok	8bc65b9318	[HUDI-1731] Rename UpsertPartitioner in hudi-java-client (#2734 ) Co-authored-by: lei.zhu <lei.zhu@envisioncn.com>	2021-03-31 11:06:04 +08:00
Gary Li	452f5e2d66	[HOTFIX] close spark session in functional test suite and disable spark3 test for spark2 (#2727 )	2021-03-29 06:04:48 -07:00
Danny Chan	d415d45416	[HUDI-1729] Asynchronous Hive sync and commits cleaning for Flink writer (#2732 )	2021-03-29 10:47:29 +08:00
Shen Hong	ecbd389a3f	[HUDI-1478] Introduce HoodieBloomIndex to hudi-java-client (#2608 )	2021-03-28 20:28:40 +08:00
n3nash	bec70413c0	[HUDI-1728] Fix MethodNotFound for HiveMetastore Locks (#2731 )	2021-03-27 10:07:10 -07:00
garyli1019	6e803e08b1	Moving to 0.9.0-SNAPSHOT on master branch.	2021-03-24 21:37:14 +08:00
n3nash	d7b18783bd	[HUDI-1709] Improving config names and adding hive metastore uri config (#2699 )	2021-03-22 01:22:06 -07:00
Jintao Guan	1277c62398	[HUDI-1653] Add support for composite keys in NonpartitionedKeyGenerator (#2627 ) * [HUDI-1653] Add support for composite keys in NonpartitionedKeyGenerator * update NonpartitionedKeyGenerator to support composite record keys * update NonpartitionedKeyGenerator	2021-03-18 15:33:31 -07:00
wangxianghu	e602e5dfb9	[MINOR] Remove unused var in AbstractHoodieWriteClient (#2693 )	2021-03-18 14:56:02 -07:00
n3nash	74241947c1	[HUDI-845] Added locking capability to allow multiple writers (#2374 ) * [HUDI-845] Added locking capability to allow multiple writers 1. Added LockProvider API for pluggable lock methodologies 2. Added Resolution Strategy API to allow for pluggable conflict resolution 3. Added TableService client API to schedule table services 4. Added Transaction Manager for wrapping actions within transactions	2021-03-16 16:43:53 -07:00
Prashant Wason	3b36cb805d	[HUDI-1552] Improve performance of key lookups from base file in Metadata Table. (#2494 ) * [HUDI-1552] Improve performance of key lookups from base file in Metadata Table. 1. Cache the KeyScanner across lookups so that the HFile index does not have to be read for each lookup. 2. Enable block caching in KeyScanner. 3. Move the lock to a limited scope of the code to reduce lock contention. 4. Removed reuse configuration * Properly close the readers, when metadata table is accessed from executors - Passing a reuse boolean into HoodieBackedTableMetadata - Preserve the fast return behavior when reusing and opening from multiple threads (no contention) - Handle concurrent close() and open readers, for reuse=false, by always synchronizing Co-authored-by: Vinoth Chandar <vinoth@apache.org>	2021-03-15 13:42:57 -07:00
Danny Chan	20786ab8a2	[HUDI-1681] Support object storage for Flink writer (#2662 ) In order to support object storage, we need these changes: * Use the Hadoop filesystem so that we can find the plugin filesystem * Do not fetch file size until the file handle is closed * Do not close the opened filesystem because we want to use the filesystem cache	2021-03-12 16:39:24 +08:00
satishkotha	c4a66324cd	[HUDI-1651] Fix archival of requested replacecommit (#2622 )	2021-03-09 15:56:44 -08:00
Raymond Xu	d3a451611c	[MINOR] HoodieClientTestHarness close resources in AfterAll phase (#2646 ) Parameterized test case like `org.apache.hudi.table.upgrade.TestUpgradeDowngrade#testUpgrade` incurs flakiness when org.apache.hadoop.fs.FileSystem#closeAll is invoked at BeforeEach; it should be invoked in AfterAll instead.	2021-03-08 17:36:03 +08:00
Shen Hong	8b9dea4ad9	[HUDI-1673] Replace scala.Tule2 to Pair in FlinkHoodieBloomIndex (#2642 )	2021-03-08 14:30:34 +08:00

1 2 3 4 5 ...

423 Commits