lanyuanxiaoyao/hudi - hudi - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
Raymond Xu	0bd38f26ca	[HUDI-2596] Make class names consistent in hudi-client (#4680 )	2022-01-27 17:05:08 -08:00
Manoj Govindassamy	f87c47352a	[HUDI-2763] Metadata table records - support for key deduplication based on hardcoded key field (#4449 ) * [HUDI-2763] Metadata table records - support for key deduplication and virtual keys - The backing log format for the metadata table is HFile, a KeyValue type. Since the key field in the metadata record payload is a duplicate of the Key in the Cell, the redundant key field in the record can be emptied to save on the cost. - HoodieHFileWriter and HoodieHFileDataBlock will now serialize records with the key field emptied by default. HFile writer tries to find if the record has metadata payload schema field 'key' and if so it does the key trimming from the record payload. - HoodieHFileReader when reading the serialized records back from disk, it materializes the missing keyFields if any. HFile reader tries to find if the record has metadata payload schema fiels 'key' and if so it does the key materialization in the record payload. - Tests have been added to verify the default virtual keys and key deduplication support for the metadata table records. Co-authored-by: Vinoth Chandar <vinoth@apache.org>	2022-01-26 13:34:04 -05:00
Alexey Kudinkin	bc7882cbe9	[HUDI-2872][HUDI-2646] Refactoring layout optimization (clustering) flow to support linear ordering (#4606 ) Refactoring layout optimization (clustering) flow to - Enable support for linear (lexicographic) ordering as one of the ordering strategies (along w/ Z-order, Hilbert) - Reconcile Layout Optimization and Clustering configuration to be more congruent	2022-01-24 16:53:54 -05:00
Sivabalan Narayanan	e00a9042e9	[HUDI-3072] Fixing conflict resolution in transaction management code path for auto commit code path (#4588 ) * Fixing conflict resolution in transaction management code path for auto commit code path * Addressing comments * Fixing test failures	2022-01-24 16:13:28 +05:30
wangxianghu	a66004a340	[HUDI-3285] Drop unused method SparkBootstrapCommitActionExecutor#handleMetadataBootstrap (#4653 )	2022-01-20 20:04:36 +04:00
YueZhang	7647562dad	[HUDI-2833][Design] Merge small archive files instead of expanding indefinitely. (#4078 ) Co-authored-by: yuezhang <yuezhang@freewheel.tv>	2022-01-18 22:42:35 -08:00
Yuwei XIAO	d36533735f	[HUDI-3194] fix MOR snapshot query during compaction (#4540 )	2022-01-17 17:24:24 -05:00
leesf	5ce45c440b	[HUDI-3172] Refactor hudi existing modules to make more code reuse in V2 Implementation (#4514 ) * Introduce hudi-spark3-common and hudi-spark2-common modules to place classes that would be reused in different spark versions, also introduce hudi-spark3.1.x to support spark 3.1.x. * Introduce hudi format under hudi-spark2, hudi-spark3, hudi-spark3.1.x modules and change the hudi format in original hudi-spark module to hudi_v1 format. * Manually tested on Spark 3.1.2 and Spark 3.2.0 SQL. * Added a README.md file under hudi-spark-datasource module.	2022-01-14 13:42:35 +08:00
董可伦	017ddbbfac	[MINOR] Fix typos (#4567 )	2022-01-11 23:17:10 -08:00
Alexey Kudinkin	f1e3762a94	[HUDI-2950] Addressing performance traps in Bulk Insert/Layout Optimization (#4234 ) * Cleaned up Z-curve/Hilbert ordering seqs: - Streamlined flow - Removed unnecessary operations (double-mapping, boxing, etc) Updated `CollectionUtils::combine` to avoid AL resizing * Tidying up * Reducing small objects churn due to Scala/Java conversions by re-using `RowFactory`, passing `Object[]` * Fixing name resolution (disambiguation overloads) * `lint` * Replaced `OverwriteAvroPayloadRecord` w/ `RewriteRecordPayload` to avoid unnecessary Avro ser/de loop * Added `PathCachingFileName` to avoid fetching substrings every time file-name is fetched; Inject `PathCachingFileName` into `HoodieWrapperFileSystem.convertPathWithScheme` * Drastically reducing size of the `ArrayDeque` allocated by `ObjectSizeCalculator` * XXX * Missing license * Fixed refs (after rebase) * Fixing compilation failure in Scala 2.11 * `PathCachingFileName` > `FileNameCachingPath` * Tidying up	2022-01-10 18:23:22 -08:00
Sivabalan Narayanan	7a8b94c82d	[HUDI-3180] Include files from completed commits while bootstrapping metadata table (#4519 )	2022-01-10 15:33:15 -05:00
Sivabalan Narayanan	56f93f4ebd	Removing rollbacks instants from timeline for restore operation (#4518 )	2022-01-10 07:44:28 +05:30
YueZhang	cf362fb2d5	[MINOR] Fix some code style issues based on check-style plugin (#4532 ) Co-authored-by: yuezhang <yuezhang@freewheel.tv>	2022-01-09 01:14:56 -08:00
Yann Byron	36790709f7	[HUDI-3125] spark-sql write timestamp directly (#4471 )	2022-01-08 23:43:25 -08:00
Sivabalan Narayanan	98ec215079	[HUDI-3178] Fixing metadata table compaction so as to not include uncommitted data (#4530 ) - There is a chance that the actual write eventually failed in data table but commit was successful in Metadata table, and if compaction was triggered in MDT, compaction could have included the uncommitted data. But once compacted, it may never be ignored while reading from metadata table. So, this patch fixes the bug. Metadata table compaction is triggered before applying the commit to metadata table to circumvent this issue.	2022-01-08 10:34:47 -05:00
Sagar Sumit	827549949c	[HUDI-2909] Handle logical type in TimestampBasedKeyGenerator (#4203 ) * [HUDI-2909] Handle logical type in TimestampBasedKeyGenerator Timestampbased key generator was returning diff values for row writer and non row writer path. this patch fixes it and is guarded by a config flag (`hoodie.datasource.write.keygenerator.consistent.logical.timestamp.enabled`)	2022-01-08 10:22:44 -05:00
Sivabalan Narayanan	8718c30324	[HUDI-3165] Enabling InProcessLockProvider for all multi-writer tests instead of FileSystemBasedLockProviderTestClass (#4427 )	2022-01-06 13:04:10 -05:00
Sivabalan Narayanan	2954027b92	[HUDI-52] Enabling savepoint and restore for MOR table (#4507 ) * Enabling restore for MOR table * Fixing savepoint for compaction commits in MOR	2022-01-06 21:26:08 +05:30
Sivabalan Narayanan	b6891d253f	[HUDI-44] Adding support to preserve commit metadata for compaction (#4428 )	2022-01-06 20:27:37 +05:30
Sagar Sumit	75133f9942	[HUDI-3170] Do not preserve filename when preserveCommitMetadata enabled (#4512 )	2022-01-05 08:09:58 -05:00
harshal	2b2ae34cb9	[HUDI-2558] Fixing Clustering w/ sort columns with null values fails (#4404 )	2022-01-03 12:19:43 +05:30
Yuwei XIAO	2444f40a4b	[HUDI-3095] abstract partition filter logic to enable code reuse (#4454 ) * [HUDI-3095] abstract partition filter logic to enable code reuse * [HUDI-3095] address reviews	2021-12-31 11:07:52 +05:30
Shawy Geng	a4e622ac61	[HUDI-1951] Add bucket hash index, compatible with the hive bucket (#3173 ) * [HUDI-2154] Add index key field to HoodieKey * [HUDI-2157] Add the bucket index and its read/write implemention of Spark engine. * revert HUDI-2154 add index key field to HoodieKey * fix all comments and introduce a new tricky way to get index key at runtime support double insert for bucket index * revert spark read optimizer based on bucket index * add the storage layout * index tag, hash function and add ut * fix ut * address partial comments * Code review feedback * add layout config and docs * fix ut * rename hoodie.layout and rebase master Co-authored-by: Vinoth Chandar <vinoth@apache.org>	2021-12-30 12:38:26 -08:00
董可伦	436becf3ea	[HUDI-2675] Fix the exception 'Not an Avro data file' when archive and clean (#4016 )	2021-12-29 22:53:17 -05:00
Yann Byron	05942e018c	[HUDI-2811] Support Spark 3.2 (#4270 )	2021-12-28 00:12:44 -08:00
Danny Chan	7b07aac286	[HUDI-3101] Excluding compaction instants from pending rollback info (#4443 )	2021-12-25 14:10:45 +08:00
Sivabalan Narayanan	1a5f8693aa	[HUDI-3011] Adding ability to read entire data with HoodieIncrSource with empty checkpoint (#4334 ) * Adding ability to read entire data with HoodieIncrSource with empty checkpoint * Addressing comments	2021-12-22 15:43:06 +05:30
Danny Chan	f1286c2c76	[HUDI-3032] Do not clean the log files right after compaction for metadata table (#4336 )	2021-12-22 11:10:27 +08:00
Raymond Xu	32a44bbe06	[HUDI-2970] Add test for archiving replace commit (#4345 )	2021-12-21 00:01:59 -05:00
Manoj Govindassamy	4a48f99a59	[HUDI-3064][HUDI-3054] FileSystemBasedLockProviderTestClass tryLock fix and TestHoodieClientMultiWriter test fixes (#4384 ) - Made FileSystemBasedLockProviderTestClass thread safe and fixed the tryLock retry logic. - Made TestHoodieClientMultiWriter. testHoodieClientBasicMultiWriter deterministic in verifying the HoodieWriteConflictException.	2021-12-19 13:31:02 -05:00
Sivabalan Narayanan	77abb5ccb9	[HUDI-3054] Fixing default lock configs for FileSystemBasedLock and fixing a flaky test (#4374 )	2021-12-18 16:15:48 -05:00
Manoj Govindassamy	d1d48ed494	[HUDI-3029] Transaction manager: avoid deadlock when doing begin and end transactions (#4363 ) * [HUDI-3029] Transaction manager: avoid deadlock when doing begin and end transactions - Transaction manager has begin and end transactions as synchronized methods. Based on the lock provider implementaion, this can lead to deadlock situation when the underlying lock() calls are blocking or with a long timeout. - Fixing transaction manager begin and end transactions to not get to deadlock and to not assume anything on the lock provider implementation.	2021-12-18 09:43:17 -05:00
xiarixiaoyao	9246b16492	[HUDI-2958] Automatically set spark.sql.parquet.writelegacyformat, when using bulkinsert to insert data which contains decimalType (#4253 )	2021-12-17 08:58:02 -05:00
xiarixiaoyao	294d712948	[HUDI-3001] Clean up the marker directory when finish bootstrap operation. (#4298 )	2021-12-16 12:36:01 -08:00
Danny Chan	ea2eba1a55	[HUDI-3015] Implement #reset and #sync for metadata filesystem view (#4307 )	2021-12-16 15:26:16 +08:00
Manoj Govindassamy	b22c2c611b	[HUDI-2938] Metadata table util to get latest file slices for reader/writers (#4218 )	2021-12-11 20:42:36 -08:00
Manoj Govindassamy	c48a2a125a	[HUDI-2527] Multi writer test with conflicting async table services (#4046 )	2021-12-10 20:01:19 -05:00
Alexey Kudinkin	2d864f7524	[HUDI-2814] Make Z-index more generic Column-Stats Index (#4106 )	2021-12-10 14:56:09 -08:00
zhangyue19921010	3ba2909690	[HUDI-2892][BUG] Pending Clustering may stain the ActiveTimeLine and lead to incomplete query results (#4172 ) Co-authored-by: yuezhang <yuezhang@freewheel.tv>	2021-12-10 09:57:01 -08:00
Sivabalan Narayanan	be368264f4	[HUDI-2952] Fixing metadata table for non-partitioned dataset (#4243 )	2021-12-10 11:11:42 -05:00
Yuwei XIAO	f194566ed4	[HUDI-2849] Improve SparkUI job description for write path (#4222 )	2021-12-10 23:22:37 +08:00
xiarixiaoyao	456d74ce4e	[HUDI-2901] Fixed the bug clustering jobs cannot running in parallel (#4178 )	2021-12-09 22:39:35 -08:00
Sivabalan Narayanan	1d4fb827e7	[HUDI-2923] Fixing metadata table reader when metadata compaction is inflight (#4206 ) * [HUDI-2923] Fixing metadata table reader when metadata compaction is inflight * Fixing retry of pending compaction in metadata table and enhancing tests	2021-12-03 21:44:50 -08:00
vinoth chandar	0fd6b2d71e	[HUDI-2933] DISABLE Metadata table by default (#4213 )	2021-12-03 21:12:35 -08:00
Raymond Xu	a799fae316	[MINOR] Mitigate CI jobs timeout issues (#4173 ) * skip shutdown zookeeper in `@AfterAll` in TestHBaseIndex * rebalance CI tests	2021-12-03 21:08:32 -08:00
Sivabalan Narayanan	e483f7c776	[HUDI-2902] Fixing populate meta fields with Hfile writers and Disabling virtual keys by default for metadata table (#4194 )	2021-12-03 07:20:21 -05:00
rmahindra123	91d2e61433	[HUDI-2904] Fix metadata table archival overstepping between regular writers and table services (#4186 ) - Co-authored-by: Rajesh Mahindra <rmahindra@Rajeshs-MacBook-Pro.local> - Co-authored-by: Sivabalan Narayanan <n.siva.b@gmail.com>	2021-12-02 13:32:26 -05:00
Alexey Kudinkin	772f5ca24e	Fixed partitions produced by layout optimization in case order-by key is composed of a single column (#4183 )	2021-12-01 20:56:04 -08:00
Shawy Geng	5284730175	[HUDI-2881] Compact the file group with larger log files to reduce write amplification (#4152 )	2021-12-02 09:41:04 +08:00
yuzhao.cyz	a1d0ff4209	Moving to 0.11.0-SNAPSHOT on master branch.	2021-11-27 17:22:10 +08:00

1 2 3 4 5 ...

270 Commits