lanyuanxiaoyao/hudi - hudi - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
xiarixiaoyao	2d73c8ae86	[HUDI-3355] Issue with out of order commits in the timeline when ingestion writers using SparkAllowUpdateStrategy (#4962 )	2022-03-30 15:54:25 -07:00
Alexey Kudinkin	8b796e9686	[HUDI-3653] Cleaning up bespoke Column Stats Index implementation (#5062 )	2022-03-30 10:01:43 -07:00
Danny Chan	b9fbada2f2	[minor] Follow 3178, fix the flink metadata table compaction (#5175 )	2022-03-30 20:45:29 +08:00
Sivabalan Narayanan	4fed8dd319	[HUDI-3485] Adding scheduler pool configs for async clustering (#5043 )	2022-03-29 21:27:45 -04:00
Alexey Kudinkin	e5a2baeed0	[HUDI-3549] Removing dependency on "spark-avro" (#4955 ) Hudi will be taking on promise for it bundles to stay compatible with Spark minor versions (for ex 2.4, 3.1, 3.2): meaning that single build of Hudi (for ex "hudi-spark3.2-bundle") will be compatible with ALL patch versions in that minor branch (in that case 3.2.1, 3.2.0, etc) To achieve that we'll have to remove (and ban) "spark-avro" as a dependency, which on a few occasions was the root-cause of incompatibility b/w consecutive Spark patch versions (most recently 3.2.1 and 3.2.0, due to this PR). Instead of bundling "spark-avro" as dependency, we will be copying over some of the classes Hudi depends on and maintain them along the Hudi code-base to make sure we're able to provide for the aforementioned guarantee. To workaround arising compatibility issues we will be applying local patches to guarantee compatibility of Hudi bundles w/in the Spark minor version branches. Following Hudi modules to Spark minor branches is currently maintained: "hudi-spark3" -> 3.2.x "hudi-spark3.1.x" -> 3.1.x "hudi-spark2" -> 2.4.x Following classes hierarchies (borrowed from "spark-avro") are maintained w/in these Spark-specific modules to guarantee compatibility with respective minor version branches: AvroSerializer AvroDeserializer AvroUtils Each of these classes has been correspondingly copied from Spark 3.2.1 (for 3.2.x branch), 3.1.2 (for 3.1.x branch), 2.4.4 (for 2.4.x branch) into their respective modules. SchemaConverters class in turn is shared across all those modules given its relative stability (there're only cosmetical changes from 2.4.4 to 3.2.1). All of the aforementioned classes have their corresponding scope of visibility limited to corresponding packages (org.apache.spark.sql.avro, org.apache.spark.sql) to make sure broader code-base does not become dependent on them and instead relies on facades abstracting them. Additionally, given that Hudi plans on supporting all the patch versions of Spark w/in aforementioned minor versions branches of Spark, additional build steps were added to validate that Hudi could be properly compiled against those versions. Testing, however, is performed against the most recent patch versions of Spark with the help of Azure CI. Brief change log: - Removing spark-avro bundling from Hudi by default - Scaffolded Spark 3.2.x hierarchy - Bootstrapped Spark 3.1.x Avro serializer/deserializer hierarchy - Bootstrapped Spark 2.4.x Avro serializer/deserializer hierarchy - Moved ExpressionCodeGen,ExpressionPayload into hudi-spark module - Fixed AvroDeserializer to stay compatible w/ both Spark 3.2.1 and 3.2.0 - Modified bot.yml to build full matrix of support Spark versions - Removed "spark-avro" dependency from all modules - Fixed relocation of spark-avro classes in bundles to assist in running integ-tests.	2022-03-29 14:44:47 -04:00
Sivabalan Narayanan	d074089c62	[HUDI-2566] Adding multi-writer test support to integ test (#5065 )	2022-03-28 17:05:00 -04:00
Y Ethan Guo	4ed84b216d	[HUDI-3720] Fix the logic of reattempting pending rollback (#5148 )	2022-03-28 14:54:31 -04:00
xiarixiaoyao	9da2dd416e	[HUDI-3719] High performance costs of AvroSerizlizer in DataSource wr… (#5137 ) * [HUDI-3719] High performance costs of AvroSerizlizer in DataSource writing * add benchmark framework which modify from spark add avroSerDerBenchmark	2022-03-27 11:01:43 -07:00
Y Ethan Guo	484b3407e0	[HUDI-3604] Adjust the order of timeline changes in rollbacks (#5114 )	2022-03-26 22:37:44 -07:00
Alexey Kudinkin	189d5297b8	[HUDI-3709] Fixing `ParquetWriter` impls not respecting Parquet Max File Size limit (#5129 )	2022-03-26 17:51:36 -04:00
RexAn	57b4f39c31	[HUDI-3612] Clustering strategy should create new TypedProperties when modifying it (#5027 )	2022-03-26 16:16:03 +05:30
Danny Chan	0c09a973fb	[HUDI-3435] Do not throw exception when instant to rollback does not exist in metadata table active timeline (#4821 )	2022-03-26 11:42:54 +08:00
Alexey Kudinkin	51034fecf1	[HUDI-3396] Refactoring `MergeOnReadRDD` to avoid duplication, fetch only projected columns (#4888 )	2022-03-25 09:32:03 -07:00
Alexey Kudinkin	8b38ddedc2	[HUDI-3594] Supporting Composite Expressions over Data Table Columns in Data Skipping flow (#4996 )	2022-03-24 22:27:15 -07:00
Danny Chan	8896864d7b	[HUDI-3678] Fix record rewrite of create handle when 'preserveMetadata' is true (#5088 )	2022-03-25 11:48:50 +08:00
Y Ethan Guo	eaa4c4f2e2	[HUDI-1180] Upgrade HBase to 2.4.9 (#5004 ) Co-authored-by: Sagar Sumit <sagarsumit09@gmail.com>	2022-03-24 19:04:53 -07:00
Danny Chan	5e86cdd1e9	[HUDI-3701] Flink bulk_insert support bucket hash index (#5118 )	2022-03-25 09:01:42 +08:00
Y Ethan Guo	608d4bf32d	[HUDI-3638] Make ZookeeperBasedLockProvider serializable (#5112 )	2022-03-24 17:59:47 -07:00
Y Ethan Guo	9b3dd2e0b7	[HUDI-3624] Check all instants before starting a commit in metadata table (#5098 )	2022-03-24 17:13:58 -07:00
Sagar Sumit	f96ba7abf0	[HUDI-3642] Handle NPE due to empty requested replacecommit metadata (#5090 )	2022-03-23 12:13:02 -07:00
Rajesh Mahindra	5f570ea151	[HUDI-2883] Refactor hive sync tool / config to use reflection and standardize configs (#4175 ) - Refactor hive sync tool / config to use reflection and standardize configs Co-authored-by: sivabalan <n.siva.b@gmail.com> Co-authored-by: Rajesh Mahindra <rmahindra@Rajeshs-MacBook-Pro.local> Co-authored-by: Raymond Xu <2701446+xushiyan@users.noreply.github.com>	2022-03-21 22:56:31 -04:00
Y Ethan Guo	9b6e138af2	[HUDI-3640] Set SimpleKeyGenerator as default in 2to3 table upgrade for Spark engine (#5075 )	2022-03-21 20:35:06 -04:00
Pratyaksh Sharma	ca0931d332	[HUDI-1436]: Provide an option to trigger clean every nth commit (#4385 ) - Provided option to trigger clean every nth commit with default number of commits as 1 so that existing users are not affected. Co-authored-by: sivabalan <n.siva.b@gmail.com>	2022-03-21 20:06:30 -04:00
wxp4532	26e5d2e6fc	[HUDI-3559] Flink bucket index with COW table throws NoSuchElementException Actually method FlinkWriteHelper#deduplicateRecords does not guarantee the records sequence, but there is a implicit constraint: all the records in one bucket should have the same bucket type(instant time here), the BucketStreamWriteFunction breaks the rule and fails to comply with this constraint. close apache/hudi#5018	2022-03-21 17:34:54 +08:00
Danny Chan	799c78e688	[HUDI-3665] Support flink multiple versions (#5072 )	2022-03-21 10:34:50 +08:00
Alexey Kudinkin	1b6e201160	[HUDI-3663] Fixing Column Stats index to properly handle first Data Table commit (#5070 ) * Fixed metadata conversion util to extract schema from `HoodieCommitMetadata` * Fixed failure to fetch columns to index in empty table * Abort indexing seq in case there are no columns to index * Fallback to index at least primary key columns, in case no writer schema could be obtained to index all columns * Fixed `getRecordFields` incorrectly ignoring default value * Make sure Hudi metadata fields are also indexed	2022-03-20 10:24:13 +05:30
Alexey Kudinkin	099c2c099a	[HUDI-3457] Refactored Spark DataSource Relations to avoid code duplication (#4877 ) Refactoring Spark DataSource Relations to avoid code duplication. Following Relations were in scope: - BaseFileOnlyViewRelation - MergeOnReadSnapshotRelaation - MergeOnReadIncrementalRelation	2022-03-18 22:32:16 -07:00
Raymond Xu	7446ff95a7	[HUDI-2439] Replace RDD with HoodieData in HoodieSparkTable and commit executors (#4856 ) - Adopt HoodieData in Spark action commit executors - Make Spark independent DeleteHelper, WriteHelper, MergeHelper in hudi-client-common - Make HoodieTable in WriteClient APIs have raw type to decouple with Client's generic types	2022-03-17 04:17:56 -07:00
Y Ethan Guo	5ba2d9ab2f	[HUDI-3494] Consider triggering condition of MOR compaction during archival (#4974 )	2022-03-17 01:28:11 -04:00
Y Ethan Guo	95e6e53810	[HUDI-3404] Automatically adjust write configs based on metadata table and write concurrency mode (#4975 )	2022-03-17 01:25:04 -04:00
Alexey Kudinkin	5e8ff8d793	[HUDI-3514] Rebase Data Skipping flow to rely on MT Column Stats index (#4948 )	2022-03-15 10:38:36 -07:00
liujinhui	e60acc1258	[HUDI-3583] Fix MarkerBasedRollbackStrategy NoSuchElementException (#4984 ) Co-authored-by: Y Ethan Guo <ethan.guoyihua@gmail.com>	2022-03-12 23:00:50 -08:00
Sivabalan Narayanan	e7bb0413af	[HUDI-3556] Re-use rollback instant for rolling back of clustering and compaction if rollback failed mid-way (#4971 )	2022-03-11 18:40:13 -05:00
Alexey Kudinkin	5d59bf67ae	[HUDI-3513] Make sure Column Stats does not fail in case it fails to load previous Index Table state (#5015 )	2022-03-11 17:39:22 -05:00
Danny Chan	ec24407191	[HUDI-3581] Reorganize some clazz for hudi flink (#4983 )	2022-03-10 15:55:15 +08:00
Alexey Kudinkin	034addaef5	[HUDI-3396] Make sure `BaseFileOnlyViewRelation` only reads projected columns (#4818 ) NOTE: This change is first part of the series to clean up Hudi's Spark DataSource related implementations, making sure there's minimal code duplication among them, implementations are consistent and performant This PR is making sure that BaseFileOnlyViewRelation only reads projected columns as well as avoiding unnecessary serde from Row to InternalRow Brief change log - Introduced HoodieBaseRDD as a base for all custom RDD impls - Extracted common fields/methods to HoodieBaseRelation - Cleaned up and streamlined HoodieBaseFileViewOnlyRelation - Fixed all of the Relations to avoid superfluous Row <> InternalRow conversions	2022-03-09 21:45:25 -05:00
MrSleeping123	8859b48b2a	[HUDI-3383] Sync column comments while syncing a hive table (#4960 ) Desc: Add a hive sync config(hoodie.datasource.hive_sync.sync_comment). This config defaults to false. While syncing data source to hudi, add column comments to source avro schema, and the sync_comment is true, syncing column comments to the hive table.	2022-03-10 09:44:39 +08:00
Sivabalan Narayanan	4324e874ae	[HUDI-3587] Making SupportsUpgradeDowngrade serializable (#4991 )	2022-03-09 00:04:42 -05:00
ForwardXu	08fd80c913	[HUDI-3221] Support querying a table as of a savepoint (#4720 )	2022-03-08 10:02:34 -08:00
Sagar Sumit	575bc63468	[HUDI-3356][HUDI-3203] HoodieData for metadata index records; BloomFilter construction from index based on the type param (#4848 ) Rework of #4761 This diff introduces following changes: - Write stats are converted to metadata index records during the commit. Making them use the HoodieData type so that the record generation scales up with needs. - Metadata index init support for bloom filter and column stats partitions. - When building the BloomFilter from the index records, using the type param stored in the payload instead of hardcoded type. - Delta writes can change column ranges and the column stats index need to be properly updated with new ranges to be consistent with the table dataset. This fix add column stats index update support for the delta writes. Co-authored-by: Manoj Govindassamy <manoj.govindassamy@gmail.com>	2022-03-08 10:39:04 -05:00
Sivabalan Narayanan	29040762fa	[HUDI-3576] Configuring timeline refreshes based on latest commit (#4973 )	2022-03-07 17:01:49 -05:00
Alexey Kudinkin	a66fd40692	[HUDI-3365] Make sure Metadata Table records are updated appropriately on HDFS (#4739 ) - This change makes sure MT records are updated appropriately on HDFS: previously after Log File append operations MT records were updated w/ just the size of the deltas being appended to the original files, which have been found to be the cause of issues in case of Rollbacks that were instead updating MT with records bearing the full file-size. - To make sure that we hedge against similar issues going f/w, this PR alleviates this discrepancy and streamlines the flow of MT table always ingesting records bearing full file-sizes.	2022-03-07 15:38:27 -05:00
Alexey Kudinkin	f0bcee3c01	[HUDI-3561] Avoid including whole `MultipleSparkJobExecutionStrategy` object into the closure for Spark to serialize (#4954 ) - Avoid including whole MultipleSparkJobExecutionStrategy object into the closure for Spark to serialize	2022-03-07 13:42:03 -05:00
Sivabalan Narayanan	3539578ccb	[HUDI-3213] Making commit preserve metadata to true for compaction (#4811 ) * Making commit preserve metadata to true * Fixing integ tests * Fixing preserve commit metadata for metadata table * fixed bootstrap tests * temp diff * Fixing merge handle * renaming fallback record * fixing build issue * Fixing test failures	2022-03-07 18:02:05 +05:30
Aditya Tiwari	051ad0b033	[HUDI-3130] Fixing Hive getSchema for RT tables addressing different partitions having different schemas (#4468 ) * Fixing Hive getSchema for RT tables * Addressing feedback * temp diff * fixing tests after spark datasource read support for metadata table is merged to master * Adding multi-partition schema evolution tests to HoodieRealTimeRecordReader Co-authored-by: Aditya Tiwari <aditya.tiwari@flipkart.com> Co-authored-by: sivabalan <n.siva.b@gmail.com>	2022-03-06 07:51:35 +05:30
Sivabalan Narayanan	6a46130037	[HUDI-2761] Fixing timeline server for repeated refreshes (#4812 ) * Fixing timeline server for repeated refreshes	2022-03-05 10:04:16 +08:00
shibei	62f534d002	[HUDI-3445] Support Clustering Command Based on Call Procedure Command for Spark SQL (#4901 ) * [HUDI-3445] Clustering Command Based on Call Procedure Command for Spark SQL * [HUDI-3445] Clustering Command Based on Call Procedure Command for Spark SQL * [HUDI-3445] Clustering Command Based on Call Procedure Command for Spark SQL Co-authored-by: shibei <huberylee.li@alibaba-inc.com>	2022-03-04 09:33:16 +08:00
Sivabalan Narayanan	876a891979	[HUDI-3544] Fixing "populate meta fields" update to metadata table (#4941 ) * Fixing populateMeta fields update to metadata table * Fix checkstyle violations Co-authored-by: Sagar Sumit <sagarsumit09@gmail.com>	2022-03-03 17:02:25 +05:30
Danny Chan	1d57bd17c2	[minor] Cosmetic changes following HUDI-3315 (#4934 )	2022-03-02 17:44:52 +08:00
Gary Li	10d866f083	[HUDI-3315] RFC-35 Part-1 Support bucket index in Flink writer (#4679 ) * Support bucket index in Flink writer * Use record key as default index key	2022-03-02 15:14:44 +08:00

1 2 3 4 5 ...

839 Commits