lanyuanxiaoyao/hudi - hudi - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
YueZhang	76e2faa28d	[HUDI-3370] The files recorded in the commit may not match the actual ones for MOR Compaction (#4753 ) * use HoodieCommitMetadata to replace writeStatuses computation Co-authored-by: yuezhang <yuezhang@freewheel.tv>	2022-02-14 11:12:52 +08:00
Y Ethan Guo	6aba00e84f	[MINOR] Fix typos in Spark client related classes (#4781 )	2022-02-13 06:41:58 -08:00
Sivabalan Narayanan	e7ec3a82dc	[HUDI-2432] Adding restore.requested instant and restore plan for restore action (#4605 ) - This adds a restore plan and serializes it to restore.requested meta file in timeline. This also means that we are introducing schedule and execution phases for restore which was not present before.	2022-02-10 08:06:23 -05:00
Sivabalan Narayanan	0ababcfaa7	[HUDI-1847] Adding inline scheduling support for spark datasource path for compaction and clustering (#4420 ) - This adds support in spark-datasource to just schedule table services inline so that users can leverage async execution w/o the need for lock service providers.	2022-02-10 08:04:55 -05:00
Y Ethan Guo	b8601a9f58	[HUDI-2656] Generalize HoodieIndex for flexible record data type (#3893 ) Co-authored-by: Raymond Xu <2701446+xushiyan@users.noreply.github.com>	2022-02-03 20:24:04 -08:00
Alexey Kudinkin	69dfcda116	[HUDI-3191] Removing duplicating file-listing process w/in Hive's MOR `FileInputFormat`s (#4556 )	2022-02-03 14:01:41 -08:00
Manoj Govindassamy	5927bdd1c0	[HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups (#4352 ) * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - Today, base files have bloom filter at their footers and index lookups have to load the base file to perform any bloom lookups. Though we have interval tree based file purging, we still end up in significant amount of base file read for the bloom filter for the end index lookups for the keys. This index lookup operation can be made more performant by having all the bloom filters in a new metadata partition and doing pointed lookups based on keys. * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - Adding indexing support for clean, restore and rollback operations. Each of these operations will now be converted to index records for bloom filter and column stats additionally. * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - Making hoodie key consistent for both column stats and bloom index by including fileId instead of fileName, in both read and write paths. - Performance optimization for looking up records in the metadata table. - Avoiding multi column sorting needed for HoodieBloomMetaIndexBatchCheckFunction * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - HoodieBloomMetaIndexBatchCheckFunction cleanup to remove unused classes - Base file checking before reading the file footer for bloom or column stats * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - Updating the bloom index and column stats index to have full file name included in the key instead of just file id. - Minor test fixes. * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - Fixed flink commit method to handle metadata table all partition update records - TestBloomIndex fixes * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - SparkHoodieBloomIndexHelper code simplification for various config modes - Signature change for getBloomFilters() and getColumnStats(). Callers can just pass in interested partition and file names, the index key is then constructed internally based on the passed in parameters. - KeyLookupHandle and KeyLookupResults code refactoring - Metadata schema changes - removed the reserved field * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - Removing HoodieColumnStatsMetadata and using HoodieColumnRangeMetadata instead. Fixed the users of the the removed class. * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - Extending meta index test to cover deletes, compactions, clean and restore table operations. Also, fixed the getBloomFilters() and getColumnStats() to account for deleted entries. * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - Addressing review comments - java doc for new classes, keys sorting for lookup, index methods renaming. * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - Consolidated the bloom filter checking for keys in to one HoodieMetadataBloomIndexCheckFunction instead of a spearate batch and lazy mode. Removed all the configs around it. - Made the metadata table partition file group count configurable. - Fixed the HoodieKeyLookupHandle to have auto closable file reader when checking bloom filter and range keys. - Config property renames. Test fixes. * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - Enabling column stats indexing for all columns by default - Handling column stat generation errors and test update * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - Metadata table partition file group count taken from the slices when the table is bootstrapped. - Prep records for the commit refactored to the base class - HoodieFileReader interface changes for filtering keys - Multi column and data types support for colums stats index * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - rebase to latest master and merge fixes for the build and test failures * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - Extending the metadata column stats type payload schema to include more statistics about the column ranges to help query integration. * [HUDI-1295] Metadata Index - Bloom filter and Column stats index to speed up index lookups - Addressing review comments	2022-02-03 18:12:48 +05:30
Alexey Kudinkin	819e8018ff	[HUDI-3322][HUDI-3343] Fixing Metadata Table Records Duplication Issues (#4716 ) This change is addressing issues in regards to Metadata Table observing ingesting duplicated records leading to it persisting incorrect file-sizes for the files referred to in those records. There are multiple issues that were leading to that: - [HUDI-3322] Incorrect Rollback Plan generation: Rollback Plan generated for MOR tables was overly expansively listing all log-files with the latest base-instant as the ones that have been affected by the rollback, leading to invalid MT records being ingested referring to those. - [HUDI-3343] Metadata Table including Uncommitted Log Files during Bootstrap: Since MT is bootstrapped at the end of the commit operation execution (after FS activity, but before committing to the timeline), it was actually incorrectly ingesting some files that were part of the intermediate state of the operation being committed. This change will unblock Stack of PRs based off #4556	2022-02-02 16:10:51 -05:00
Raymond Xu	0bd38f26ca	[HUDI-2596] Make class names consistent in hudi-client (#4680 )	2022-01-27 17:05:08 -08:00
Manoj Govindassamy	f87c47352a	[HUDI-2763] Metadata table records - support for key deduplication based on hardcoded key field (#4449 ) * [HUDI-2763] Metadata table records - support for key deduplication and virtual keys - The backing log format for the metadata table is HFile, a KeyValue type. Since the key field in the metadata record payload is a duplicate of the Key in the Cell, the redundant key field in the record can be emptied to save on the cost. - HoodieHFileWriter and HoodieHFileDataBlock will now serialize records with the key field emptied by default. HFile writer tries to find if the record has metadata payload schema field 'key' and if so it does the key trimming from the record payload. - HoodieHFileReader when reading the serialized records back from disk, it materializes the missing keyFields if any. HFile reader tries to find if the record has metadata payload schema fiels 'key' and if so it does the key materialization in the record payload. - Tests have been added to verify the default virtual keys and key deduplication support for the metadata table records. Co-authored-by: Vinoth Chandar <vinoth@apache.org>	2022-01-26 13:34:04 -05:00
Sivabalan Narayanan	e00a9042e9	[HUDI-3072] Fixing conflict resolution in transaction management code path for auto commit code path (#4588 ) * Fixing conflict resolution in transaction management code path for auto commit code path * Addressing comments * Fixing test failures	2022-01-24 16:13:28 +05:30
YueZhang	7647562dad	[HUDI-2833][Design] Merge small archive files instead of expanding indefinitely. (#4078 ) Co-authored-by: yuezhang <yuezhang@freewheel.tv>	2022-01-18 22:42:35 -08:00
Yuwei XIAO	d36533735f	[HUDI-3194] fix MOR snapshot query during compaction (#4540 )	2022-01-17 17:24:24 -05:00
Sivabalan Narayanan	7a8b94c82d	[HUDI-3180] Include files from completed commits while bootstrapping metadata table (#4519 )	2022-01-10 15:33:15 -05:00
Sivabalan Narayanan	56f93f4ebd	Removing rollbacks instants from timeline for restore operation (#4518 )	2022-01-10 07:44:28 +05:30
Yann Byron	36790709f7	[HUDI-3125] spark-sql write timestamp directly (#4471 )	2022-01-08 23:43:25 -08:00
Sivabalan Narayanan	98ec215079	[HUDI-3178] Fixing metadata table compaction so as to not include uncommitted data (#4530 ) - There is a chance that the actual write eventually failed in data table but commit was successful in Metadata table, and if compaction was triggered in MDT, compaction could have included the uncommitted data. But once compacted, it may never be ignored while reading from metadata table. So, this patch fixes the bug. Metadata table compaction is triggered before applying the commit to metadata table to circumvent this issue.	2022-01-08 10:34:47 -05:00
Sagar Sumit	827549949c	[HUDI-2909] Handle logical type in TimestampBasedKeyGenerator (#4203 ) * [HUDI-2909] Handle logical type in TimestampBasedKeyGenerator Timestampbased key generator was returning diff values for row writer and non row writer path. this patch fixes it and is guarded by a config flag (`hoodie.datasource.write.keygenerator.consistent.logical.timestamp.enabled`)	2022-01-08 10:22:44 -05:00
Sivabalan Narayanan	8718c30324	[HUDI-3165] Enabling InProcessLockProvider for all multi-writer tests instead of FileSystemBasedLockProviderTestClass (#4427 )	2022-01-06 13:04:10 -05:00
Sivabalan Narayanan	2954027b92	[HUDI-52] Enabling savepoint and restore for MOR table (#4507 ) * Enabling restore for MOR table * Fixing savepoint for compaction commits in MOR	2022-01-06 21:26:08 +05:30
Sivabalan Narayanan	b6891d253f	[HUDI-44] Adding support to preserve commit metadata for compaction (#4428 )	2022-01-06 20:27:37 +05:30
Sagar Sumit	75133f9942	[HUDI-3170] Do not preserve filename when preserveCommitMetadata enabled (#4512 )	2022-01-05 08:09:58 -05:00
Yuwei XIAO	2444f40a4b	[HUDI-3095] abstract partition filter logic to enable code reuse (#4454 ) * [HUDI-3095] abstract partition filter logic to enable code reuse * [HUDI-3095] address reviews	2021-12-31 11:07:52 +05:30
Shawy Geng	a4e622ac61	[HUDI-1951] Add bucket hash index, compatible with the hive bucket (#3173 ) * [HUDI-2154] Add index key field to HoodieKey * [HUDI-2157] Add the bucket index and its read/write implemention of Spark engine. * revert HUDI-2154 add index key field to HoodieKey * fix all comments and introduce a new tricky way to get index key at runtime support double insert for bucket index * revert spark read optimizer based on bucket index * add the storage layout * index tag, hash function and add ut * fix ut * address partial comments * Code review feedback * add layout config and docs * fix ut * rename hoodie.layout and rebase master Co-authored-by: Vinoth Chandar <vinoth@apache.org>	2021-12-30 12:38:26 -08:00
董可伦	436becf3ea	[HUDI-2675] Fix the exception 'Not an Avro data file' when archive and clean (#4016 )	2021-12-29 22:53:17 -05:00
Danny Chan	7b07aac286	[HUDI-3101] Excluding compaction instants from pending rollback info (#4443 )	2021-12-25 14:10:45 +08:00
Sivabalan Narayanan	1a5f8693aa	[HUDI-3011] Adding ability to read entire data with HoodieIncrSource with empty checkpoint (#4334 ) * Adding ability to read entire data with HoodieIncrSource with empty checkpoint * Addressing comments	2021-12-22 15:43:06 +05:30
Raymond Xu	32a44bbe06	[HUDI-2970] Add test for archiving replace commit (#4345 )	2021-12-21 00:01:59 -05:00
Manoj Govindassamy	4a48f99a59	[HUDI-3064][HUDI-3054] FileSystemBasedLockProviderTestClass tryLock fix and TestHoodieClientMultiWriter test fixes (#4384 ) - Made FileSystemBasedLockProviderTestClass thread safe and fixed the tryLock retry logic. - Made TestHoodieClientMultiWriter. testHoodieClientBasicMultiWriter deterministic in verifying the HoodieWriteConflictException.	2021-12-19 13:31:02 -05:00
Sivabalan Narayanan	77abb5ccb9	[HUDI-3054] Fixing default lock configs for FileSystemBasedLock and fixing a flaky test (#4374 )	2021-12-18 16:15:48 -05:00
Danny Chan	ea2eba1a55	[HUDI-3015] Implement #reset and #sync for metadata filesystem view (#4307 )	2021-12-16 15:26:16 +08:00
Manoj Govindassamy	b22c2c611b	[HUDI-2938] Metadata table util to get latest file slices for reader/writers (#4218 )	2021-12-11 20:42:36 -08:00
Manoj Govindassamy	c48a2a125a	[HUDI-2527] Multi writer test with conflicting async table services (#4046 )	2021-12-10 20:01:19 -05:00
Alexey Kudinkin	2d864f7524	[HUDI-2814] Make Z-index more generic Column-Stats Index (#4106 )	2021-12-10 14:56:09 -08:00
zhangyue19921010	3ba2909690	[HUDI-2892][BUG] Pending Clustering may stain the ActiveTimeLine and lead to incomplete query results (#4172 ) Co-authored-by: yuezhang <yuezhang@freewheel.tv>	2021-12-10 09:57:01 -08:00
Sivabalan Narayanan	be368264f4	[HUDI-2952] Fixing metadata table for non-partitioned dataset (#4243 )	2021-12-10 11:11:42 -05:00
Sivabalan Narayanan	1d4fb827e7	[HUDI-2923] Fixing metadata table reader when metadata compaction is inflight (#4206 ) * [HUDI-2923] Fixing metadata table reader when metadata compaction is inflight * Fixing retry of pending compaction in metadata table and enhancing tests	2021-12-03 21:44:50 -08:00
vinoth chandar	0fd6b2d71e	[HUDI-2933] DISABLE Metadata table by default (#4213 )	2021-12-03 21:12:35 -08:00
Raymond Xu	a799fae316	[MINOR] Mitigate CI jobs timeout issues (#4173 ) * skip shutdown zookeeper in `@AfterAll` in TestHBaseIndex * rebalance CI tests	2021-12-03 21:08:32 -08:00
Sivabalan Narayanan	e483f7c776	[HUDI-2902] Fixing populate meta fields with Hfile writers and Disabling virtual keys by default for metadata table (#4194 )	2021-12-03 07:20:21 -05:00
rmahindra123	91d2e61433	[HUDI-2904] Fix metadata table archival overstepping between regular writers and table services (#4186 ) - Co-authored-by: Rajesh Mahindra <rmahindra@Rajeshs-MacBook-Pro.local> - Co-authored-by: Sivabalan Narayanan <n.siva.b@gmail.com>	2021-12-02 13:32:26 -05:00
Shawy Geng	5284730175	[HUDI-2881] Compact the file group with larger log files to reduce write amplification (#4152 )	2021-12-02 09:41:04 +08:00
Manoj Govindassamy	2c7656c35f	[HUDI-2475] [HUDI-2862] Metadata table creation and avoid bootstrapping race for write client & add locking for upgrade (#4114 ) Co-authored-by: Sivabalan Narayanan <n.siva.b@gmail.com>	2021-11-26 23:19:26 -08:00
Y Ethan Guo	d1e83e4ba0	[HUDI-2767] Enabling timeline-server-based marker as default (#4112 ) - Changes the default config of marker type (HoodieWriteConfig.MARKERS_TYPE or hoodie.write.markers.type) from DIRECT to TIMELINE_SERVER_BASED for Spark Engine. - Adds engine-specific marker type configs: Spark -> TIMELINE_SERVER_BASED, Flink -> DIRECT, Java -> DIRECT. - Uses DIRECT markers as well for Spark structured streaming due to timeline server only available for the first mini-batch. - Fixes the marker creation method for non-partitioned table in TimelineServerBasedWriteMarkers. - Adds the fallback to direct markers even when TIMELINE_SERVER_BASED is configured, in WriteMarkersFactory: when HDFS is used, or embedded timeline server is disabled, the fallback to direct markers happens. - Fixes the closing of timeline service. - Fixes tests that depend on markers, mainly by starting the timeline service for each test.	2021-11-26 16:41:05 -05:00
Sivabalan Narayanan	f8e0176eb0	[HUDI-2861] Re-use same rollback instant time for failed rollbacks (#4123 )	2021-11-26 16:36:42 -05:00
Sivabalan Narayanan	a88691fed3	[MINOR] Fixing test failure to fix CI build failure (#4132 )	2021-11-26 13:50:10 -05:00
Alexey Kudinkin	5755ff25a4	[HUDI-2814] Addressing issues w/ Z-order Layout Optimization (#4060 ) * `ZCurveOptimizeHelper` > `ZOrderingIndexHelper`; Moved Z-index helper under `hudi.index.zorder` package * Tidying up `ZOrderingIndexHelper` * Fixing compilation * Fixed index new/original table merging sequence to always prefer values from new index; Cleaned up `HoodieSparkUtils` * Added test for `mergeIndexSql` * Abstracted Z-index name composition w/in `ZOrderingIndexHelper`; * Fixed `DataSkippingUtils` to interrupt prunning in case data filter contains non-indexed column reference * Properly handle exceptions origination during pruning in `HoodieFileIndex` * Make sure no errors are logged upon encountering `AnalysisException` * Cleaned up Z-index updating sequence; Tidying up comments, java-docs; * Fixed Z-index to properly handle changes of the list of clustered columns * Tidying up * `lint` * Suppressing `JavaDocStyle` first sentence check * Fixed compilation * Fixing incorrect `DecimalType` conversion * Refactored test `TestTableLayoutOptimization` - Added Z-index table composition test (against fixtures) - Separated out GC test; Tidying up * Fixed tests re-shuffling column order for Z-Index table `DataFrame` to align w/ the one by one loaded from JSON * Scaffolded `DataTypeUtils` to do basic checks of Spark types; Added proper compatibility checking b/w old/new index-tables * Added test for Z-index tables merging * Fixed import being shaded by creating internal `hudi.util` package * Fixed packaging for `TestOptimizeTable` * Revised `updateMetadataIndex` seq to provide Z-index updating process w/ source table schema * Make sure existing Z-index table schema is sync'd to source table's one * Fixed shaded refs * Fixed tests * Fixed type conversion of Parquet provided metadata values into Spark expected schemas * Fixed `composeIndexSchema` utility to propose proper schema * Added more tests for Z-index: - Checking that Z-index table is built correctly - Checking that Z-index tables are merged correctly (during update) * Fixing source table * Fixing tests to read from Parquet w/ proper schema * Refactored `ParquetUtils` utility reading stats from Parquet footers * Fixed incorrect handling of Decimals extracted from Parquet footers * Worked around issues in javac failign to compile stream's collection * Fixed handling of `Date` type * Fixed handling of `DateType` to be parsed as `LocalDate` * Updated fixture; Make sure test loads Z-index fixture using proper schema * Removed superfluous scheme adjusting when reading from Parquet, since Spark is actually able to perfectly restore schema (given Parquet was previously written by Spark as well) * Fixing race-condition in Parquet's `DateStringifier` trying to share `SimpleDataFormat` object which is inherently not thread-safe * Tidying up * Make sure schema is used upon reading to validate input files are in the appropriate format; Tidying up; * Worked around javac (1.8) inability to infer expression type properly * Updated fixtures; Tidying up * Fixing compilation after rebase * Assert clustering have in Z-order layout optimization testing * Tidying up exception messages * XXX * Added test validating Z-index lookup filter correctness * Added more test-cases; Tidying up * Added tests for string expressions * Fixed incorrect Z-index filter lookup translations * Added more test-cases * Added proper handling on complex negations of AND/OR expressions by pushing NOT operator down into inner expressions for appropriate handling * Added `-target:jvm-1.8` for `hudi-spark` module * Adding more tests * Added tests for non-indexed columns * Properly handle non-indexed columns by falling back to a re-write of containing expression as `TrueLiteral` instead * Fixed tests * Removing the parquet test files and disabling corresponding tests Co-authored-by: Vinoth Chandar <vinoth@apache.org>	2021-11-26 10:02:15 -08:00
mincwang	e554c7f468	[HUDI-2852] Table metadata returns empty for non-exist partition (#4117 ) * [HUDI-2852] Table metadata returns empty for non-exist partition * add unit test * fix code checkstyle Co-authored-by: wangminchao <wangminchao@asinking.com>	2021-11-26 16:24:03 +08:00
Sivabalan Narayanan	8e1379384a	[HUDI-2841] Fixing lazy rollback for MOR with list based strategy (#4110 )	2021-11-25 16:06:04 -05:00
Sivabalan Narayanan	a9bd20804b	[HUDI-2792] Configure metadata payload consistency check (#4035 ) - Relax metadata payload consistency check to consider spark task failures with spurious deletes	2021-11-24 21:56:31 -05:00

1 2 3 4

199 Commits