lanyuanxiaoyao/hudi - hudi - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
hiscat	d035fcbb3c	[HUDI-1767] Add setter to HoodieKey and HoodieRecordLocation to have better SE/DE performance for Flink (#2779 )	2021-04-07 14:13:31 +08:00
li36909	8527590772	[HUDI-1750] Fail to load user's class if user move hudi-spark-bundle jar into spark classpath (#2753 )	2021-04-06 22:33:32 -04:00
Harshit Mittal	e692c704da	[MINOR] Fix deprecated build link for travis (#2778 )	2021-04-07 08:57:10 +08:00
Danny Chan	9c369c607d	[HUDI-1757] Assigns the buckets by record key for Flink writer (#2757 ) Currently we assign the buckets by record partition path which could cause hotspot if the partition field is datetime type. Changes to assign buckets by grouping the record whth their key first, the assignment is valid if only there is no conflict(two task write to the same bucket). This patch also changes the coordinator execution to be asynchronous.	2021-04-06 19:06:41 +08:00
li36909	920537cac8	[HUDI-1749] Clean/Compaction/Rollback command maybe never exit when operation fail (#2752 )	2021-04-05 23:23:15 -07:00
Harshit Mittal	e970e1f483	[HUDI-1696] add apache commons-codec dependency to flink-bundle explicitly (#2758 )	2021-04-01 23:07:30 -07:00
Roc Marshal	94a5e72f16	[HUDI-1737][hudi-client] Code Cleanup: Extract common method in HoodieCreateHandle & FlinkCreateHandle (#2745 )	2021-04-02 11:39:05 +08:00
pengzhiwei	684622c7c9	[HUDI-1591] Implement Spark's FileIndex for Hudi to support queries via Hudi DataSource using non-globbed table path and partition pruning (#2651 )	2021-04-01 11:12:28 -07:00
Danny Chan	9804662bc8	[HUDI-1738] Emit deletes for flink MOR table streaming read (#2742 ) Current we did a soft delete for DELETE row data when writes into hoodie table. For streaming read of MOR table, the Flink reader detects the delete records and still emit them if the record key semantics are still kept. This is useful and actually a must for streaming ETL pipeline incremental computation.	2021-04-01 15:25:31 +08:00
vinoyang	fe16d0de7c	[MINOR] Delete useless UpsertPartitioner for flink integration (#2746 )	2021-03-31 16:36:42 +08:00
Sebastian Bernauer	aa0da72c59	Preparation for Avro update (#2650 )	2021-03-30 21:50:17 -07:00
leo-Iamok	8bc65b9318	[HUDI-1731] Rename UpsertPartitioner in hudi-java-client (#2734 ) Co-authored-by: lei.zhu <lei.zhu@envisioncn.com>	2021-03-31 11:06:04 +08:00
vinoyang	3cab928b50	[HUDI-1735] Add hive-exec dependency for hudi-examples (#2737 )	2021-03-30 21:35:16 +08:00
Gary Li	050626ad6c	[MINOR] Add Missing Apache License to test files (#2736 )	2021-03-29 07:17:23 -07:00
garyli1019	e069b64e10	[HOTFIX] fix deploy staging jars script	2021-03-29 06:04:48 -07:00
Gary Li	4db970dc8a	[HOTFIX] Disable ITs for Spark3 and scala2.12 (#2733 )	2021-03-29 06:04:48 -07:00
Gary Li	452f5e2d66	[HOTFIX] close spark session in functional test suite and disable spark3 test for spark2 (#2727 )	2021-03-29 06:04:48 -07:00
Danny Chan	d415d45416	[HUDI-1729] Asynchronous Hive sync and commits cleaning for Flink writer (#2732 )	2021-03-29 10:47:29 +08:00
Shen Hong	ecbd389a3f	[HUDI-1478] Introduce HoodieBloomIndex to hudi-java-client (#2608 )	2021-03-28 20:28:40 +08:00
n3nash	bec70413c0	[HUDI-1728] Fix MethodNotFound for HiveMetastore Locks (#2731 )	2021-03-27 10:07:10 -07:00
Danny Chan	8b774fe331	[HUDI-1495] Bump Flink version to 1.12.2 (#2718 )	2021-03-26 14:25:57 +08:00
garyli1019	6e803e08b1	Moving to 0.9.0-SNAPSHOT on master branch.	2021-03-24 21:37:14 +08:00
Danny Chan	29b79c99b0	[hotfix] Log the error message for creating table source first (#2711 )	2021-03-24 18:25:37 +08:00
n3nash	01a1d7997b	[HUDI-1712] Rename & standardize config to match other configs (#2708 )	2021-03-24 17:24:02 +08:00
Danny Chan	03668dbaf1	[HUDI-1710] Read optimized query type for Flink batch reader (#2702 ) Read optimized query returns the records from: * COW table: the latest parquet files * MOR table: parquet file records from the latest compaction committed	2021-03-23 18:41:30 -07:00
legendtkl	0e6909d3e2	[MINOR][DOCUMENT] Update README doc for integ test (#2703 )	2021-03-23 20:21:56 +08:00
n3nash	d7b18783bd	[HUDI-1709] Improving config names and adding hive metastore uri config (#2699 )	2021-03-22 01:22:06 -07:00
Liulietong	ce3e8ec870	[HUDI-1667]: Fix a null value related bug for spark vectorized reader. (#2636 )	2021-03-20 07:54:20 -07:00
Volodymyr Burenin	900de34e45	[HUDI-1650] Custom avro kafka deserializer. (#2619 ) * Custom avro kafka deserializer Co-authored-by: volodymyr.burenin <volodymyr.burenin@cloudkitchens.com> Co-authored-by: Sivabalan Narayanan <sivabala@uber.com>	2021-03-20 00:51:08 -07:00
Sivabalan Narayanan	161d530f93	Fixing kafka auto.reset.offsets config param key (#2691 )	2021-03-19 12:54:29 -07:00
Sivabalan Narayanan	55a489c769	[1568] Fixing spark3 bundles (#2625 ) - [1568] Fixing spark3 bundles	2021-03-19 14:21:36 -04:00
Danny Chan	f74828fca1	[HUDI-1705] Flush as per data bucket for mini-batch write (#2695 ) Detects the buffer size for each data bucket before flushing. So that we avoid flushing data buckets with few records.	2021-03-19 16:30:54 +08:00
Jintao Guan	1277c62398	[HUDI-1653] Add support for composite keys in NonpartitionedKeyGenerator (#2627 ) * [HUDI-1653] Add support for composite keys in NonpartitionedKeyGenerator * update NonpartitionedKeyGenerator to support composite record keys * update NonpartitionedKeyGenerator	2021-03-18 15:33:31 -07:00
wangxianghu	e602e5dfb9	[MINOR] Remove unused var in AbstractHoodieWriteClient (#2693 )	2021-03-18 14:56:02 -07:00
xiarixiaoyao	d429169ff7	[HUDI-1688]hudi write should uncache rdd， when the write operation is finnished (#2673 )	2021-03-18 10:19:18 -07:00
Danny Chan	f1e0018f12	[HUDI-1704] Use PRIMARY KEY syntax to define record keys for Flink Hudi table (#2694 ) The SQL PRIMARY KEY semantics is very same with Hoodie record key, using PRIMARY KEY is more straight-forward way instead of a table option: hoodie.datasource.write.recordkey.field. After this change, both PRIMARY KEY and table option can define hoodie record key, while the PRIMARY KEY has higher priority if both are defined. Note: a column with PRIMARY KEY constraint is forced to be non-nullable.	2021-03-18 20:21:52 +08:00
Danny Chan	968488fa3a	[HUDI-1701] Implement HoodieTableSource.explainSource for all kinds of pushing down (#2690 ) We should implement the interface HoodieTableSource.explainSource to track the table source signature diff for all kinds of pushing down, such as filter pushing or limit pushing.	2021-03-17 23:05:18 +08:00
n3nash	74241947c1	[HUDI-845] Added locking capability to allow multiple writers (#2374 ) * [HUDI-845] Added locking capability to allow multiple writers 1. Added LockProvider API for pluggable lock methodologies 2. Added Resolution Strategy API to allow for pluggable conflict resolution 3. Added TableService client API to schedule table services 4. Added Transaction Manager for wrapping actions within transactions	2021-03-16 16:43:53 -07:00
Sivabalan Narayanan	b038623ed3	[HUDI 1615] Fixing null schema in bulk_insert row writer path (#2653 ) * [HUDI-1615] Avoid passing in null schema from row writing/deltastreamer * Fixing null schema in bulk insert row writer path * Fixing tests Co-authored-by: vc <vinoth@apache.org>	2021-03-16 09:44:11 -07:00
Vinoth Govindarajan	16864aee14	[HUDI-1695] Fixed the error messaging (#2679 )	2021-03-16 11:30:26 +08:00
Prashant Wason	3b36cb805d	[HUDI-1552] Improve performance of key lookups from base file in Metadata Table. (#2494 ) * [HUDI-1552] Improve performance of key lookups from base file in Metadata Table. 1. Cache the KeyScanner across lookups so that the HFile index does not have to be read for each lookup. 2. Enable block caching in KeyScanner. 3. Move the lock to a limited scope of the code to reduce lock contention. 4. Removed reuse configuration * Properly close the readers, when metadata table is accessed from executors - Passing a reuse boolean into HoodieBackedTableMetadata - Preserve the fast return behavior when reusing and opening from multiple threads (no contention) - Handle concurrent close() and open readers, for reuse=false, by always synchronizing Co-authored-by: Vinoth Chandar <vinoth@apache.org>	2021-03-15 13:42:57 -07:00
Danny Chan	76bf2cc790	[HUDI-1692] Bounded source for stream writer (#2674 ) Supports bounded source such as VALUES for stream mode writer.	2021-03-15 19:42:36 +08:00
Danny Chan	fc6c5f4285	[HUDI-1684] Tweak hudi-flink-bundle module pom and reorganize the pacakges for hudi-flink module (#2669 ) * Add required dependencies for hudi-flink-bundle module * Some packages reorganization of hudi-flink module	2021-03-15 16:02:05 +08:00
Sivabalan Narayanan	e93c6a5693	[HUDI-1496] Fixing input stream detection of GCS FileSystem (#2500 ) * Adding SchemeAwareFSDataInputStream for abstract out special handling for GCSFileSystem * Moving wrapping of fsDataInputStream to separate method in HoodieLogFileReader Co-authored-by: Vinoth Chandar <vinoth@apache.org>	2021-03-14 00:57:57 -08:00
Ankush Kanungo	f5e31be086	[HUDI-1685] keep updating current date for every batch (#2671 )	2021-03-12 15:53:01 -08:00
Danny Chan	20786ab8a2	[HUDI-1681] Support object storage for Flink writer (#2662 ) In order to support object storage, we need these changes: * Use the Hadoop filesystem so that we can find the plugin filesystem * Do not fetch file size until the file handle is closed * Do not close the opened filesystem because we want to use the filesystem cache	2021-03-12 16:39:24 +08:00
Danny Chan	e8e6708aea	[HUDI-1664] Avro schema inference for Flink SQL table (#2658 ) A Flink SQL table has DDL that defines the table schema, we can use that to infer the Avro schema and there is no need to declare a Avro schema explicitly anymore. But we still keep the config option for explicit Avro schema in case there is corner cases that the inferred schema is not correct (especially for the nullability).	2021-03-11 19:45:48 +08:00
Danny Chan	12ff562d2b	[HUDI-1678] Row level delete for Flink sink (#2659 )	2021-03-11 19:44:06 +08:00
Danny Chan	2fdae6835c	[HUDI-1663] Streaming read for Flink MOR table (#2640 ) Supports two read modes: * Read the full data set starting from the latest commit instant and subsequent incremental data set * Read data set that starts from a specified commit instant	2021-03-10 22:44:06 +08:00
satishkotha	c4a66324cd	[HUDI-1651] Fix archival of requested replacecommit (#2622 )	2021-03-09 15:56:44 -08:00

1 2 3 4 5 ...

1445 Commits