lanyuanxiaoyao/hudi - hudi - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
yuzhao.cyz	a1d0ff4209	Moving to 0.11.0-SNAPSHOT on master branch.	2021-11-27 17:22:10 +08:00
wenningd	1ee12cfa6f	[HUDI-2314] Add support for DynamoDb based lock provider (#3486 ) - Co-authored-by: Wenning Ding <wenningd@amazon.com> - Co-authored-by: Sivabalan Narayanan <n.siva.b@gmail.com>	2021-11-17 12:09:31 -05:00
Alexey Kudinkin	cbcbec4d38	[MINOR] Fixed checkstyle config to be based off Maven root-dir (requires Maven >=3.3.1 to work properly); (#4009 ) Updated README	2021-11-16 21:30:16 -05:00
Yann Byron	1f17467f73	[HUDI-1869] Upgrading Spark3 To 3.1 (#3844 ) Co-authored-by: pengzhiwei <pengzhiwei2015@icloud.com>	2021-11-02 18:25:12 -07:00
Sivabalan Narayanan	f9bc3e03e5	[MINOR] Adding a deprecated constructor to AbstractSyncHoodieClient (#3902 )	2021-11-02 12:16:38 -04:00
Sagar Sumit	5302b9a4ef	[HUDI-2662] Downloads from Nexus Pentaho repo taking too long (#3901 ) Co-authored-by: Sivabalan Narayanan <n.siva.b@gmail.com>	2021-11-01 19:14:48 -04:00
vinoyang	b1c4acf0ae	[HUDI-2614] Remove duplicated hadoop-hdfs with tests classifier exists in bundles (#3864 )	2021-10-26 22:36:10 +08:00
rmahindra123	3686c25fae	[HUDI-2469] [Kafka Connect] Replace json based payload with protobuf for Transaction protocol. (#3694 ) * Substitue Control Event with protobuf * Fix tests * Fix unit tests * Add javadocs * Add javadocs * Address reviewer comments Co-authored-by: Rajesh Mahindra <rmahindra@Rajeshs-MacBook-Pro.local>	2021-10-19 14:29:48 -07:00
rmahindra123	e528dd798a	[HUDI-2394] Implement Kafka Sink Protocol for Hudi for Ingesting Immutable Data (#3592 ) - Fixing packaging, naming of classes - Use of log4j over slf4j for uniformity - More follow-on fixes - Added a version to control/coordinator events. - Eliminated the config added to write config - Fixed fetching of checkpoints based on table type - Clean up of naming, code placement Co-authored-by: Rajesh Mahindra <rmahindra@Rajeshs-MacBook-Pro.local> Co-authored-by: Vinoth Chandar <vinoth@apache.org>	2021-09-10 18:20:26 -07:00
Raymond Xu	38c9b85aa8	[HUDI-2280] Use GitHub Actions to build different scala spark versions (#3556 )	2021-09-01 08:51:00 -07:00
Danny Chan	66f951322a	[HUDI-2191] Bump flink version to 1.13.1 (#3291 )	2021-08-16 18:14:05 +08:00
Udit Mehrotra	3e301196bf	Moving to 0.10.0-SNAPSHOT on master branch.	2021-08-14 18:51:09 -07:00
Sagar Sumit	5cc96e85c1	[HUDI-1897] Deltastreamer source for AWS S3 (#3433 ) - Added two sources for two stage pipeline. a. S3EventsSource that fetches events from SQS and ingests to a meta hoodie table. b. S3EventsHoodieIncrSource reads S3 events from this meta hoodie table, fetches actual objects from S3 and ingests to sink hoodie table. - Added selectors to assist in S3EventsSource. Co-authored-by: Satish M <84978833+satishmittal1111@users.noreply.github.com> Co-authored-by: Vinoth Chandar <vinoth@apache.org>	2021-08-14 08:25:10 -04:00
pengzhiwei	3f8ca1a355	[HUDI-2182] Support Compaction Command For Spark Sql (#3277 )	2021-08-06 15:12:10 +08:00
pengzhiwei	0dcd6a8fca	[HUDI-2233] Use HMS To Sync Hive Meta For Spark Sql (#3387 )	2021-08-05 09:57:22 -04:00
pengzhiwei	151f22e43a	[HUDI-2195] Sync Hive Failed When Execute CTAS In Spark2 And Spark3 (#3299 )	2021-07-22 15:33:38 +08:00
Vinay Patil	5a94b6bf54	[HUDI-2192] Clean up Multiple versions of scala libraries detected Warning (#3292 )	2021-07-21 00:33:27 -07:00
Randal Boyle	60e0254e67	[HUDI-1996] Adding functionality to allow the providing of basic auth creds for confluent cloud schema registry (#3097 ) * adding support for basic auth with confluent cloud schema registry	2021-07-05 23:40:23 -07:00
Jintao Guan	b8fe5b91d5	[HUDI-764] [HUDI-765] ORC reader writer Implementation (#2999 ) Co-authored-by: Qingyun (Teresa) Kang <kteresa@uber.com>	2021-06-15 15:21:43 -07:00
Raymond Xu	f922837064	[HUDI-1950] Fix Azure CI failure in TestParquetUtils (#2984 ) * fix azure pipeline configs * add pentaho.org in maven repositories * Make sure file paths with scheme in TestParquetUtils * add azure build status to README	2021-06-15 03:45:17 -07:00
pengzhiwei	f760ec543e	[HUDI-1659] Basic Implement Of Spark Sql Support For Hoodie (#2645 ) Main functions: Support create table for hoodie. Support CTAS. Support Insert for hoodie. Including dynamic partition and static partition insert. Support MergeInto for hoodie. Support DELETE Support UPDATE Both support spark2 & spark3 based on DataSourceV1. Main changes: Add sql parser for spark2. Add HoodieAnalysis for sql resolve and logical plan rewrite. Add commands implementation for CREATE TABLE、INSERT、MERGE INTO & CTAS. In order to push down the update&insert logical to the HoodieRecordPayload for MergeInto, I make same change to the HoodieWriteHandler and other related classes. 1、Add the inputSchema for parser the incoming record. This is because the inputSchema for MergeInto is different from writeSchema as there are some transforms in the update& insert expression. 2、Add WRITE_SCHEMA to HoodieWriteConfig to pass the write schema for merge into. 3、Pass properties to HoodieRecordPayload#getInsertValue to pass the insert expression and table schema. Verify this pull request Add TestCreateTable for test create hoodie tables and CTAS. Add TestInsertTable for test insert hoodie tables. Add TestMergeIntoTable for test merge hoodie tables. Add TestUpdateTable for test update hoodie tables. Add TestDeleteTable for test delete hoodie tables. Add TestSqlStatement for test supported ddl/dml currently.	2021-06-07 23:24:32 -07:00
vinoth chandar	d02c0e5387	[MINOR] Resolve build issue arising from inaccessible pentaho jar (#3034 ) - Fixes #160 #2479	2021-06-04 15:28:44 -04:00
Raymond Xu	3418a92de8	[HUDI-1620] Fix Metrics UT (#2894 ) Make sure shutdown Metrics between unit test cases to ensure isolation	2021-04-30 11:20:41 -07:00
Gary Li	4db970dc8a	[HOTFIX] Disable ITs for Spark3 and scala2.12 (#2733 )	2021-03-29 06:04:48 -07:00
Gary Li	452f5e2d66	[HOTFIX] close spark session in functional test suite and disable spark3 test for spark2 (#2727 )	2021-03-29 06:04:48 -07:00
Danny Chan	8b774fe331	[HUDI-1495] Bump Flink version to 1.12.2 (#2718 )	2021-03-26 14:25:57 +08:00
garyli1019	6e803e08b1	Moving to 0.9.0-SNAPSHOT on master branch.	2021-03-24 21:37:14 +08:00
Sivabalan Narayanan	55a489c769	[1568] Fixing spark3 bundles (#2625 ) - [1568] Fixing spark3 bundles	2021-03-19 14:21:36 -04:00
n3nash	74241947c1	[HUDI-845] Added locking capability to allow multiple writers (#2374 ) * [HUDI-845] Added locking capability to allow multiple writers 1. Added LockProvider API for pluggable lock methodologies 2. Added Resolution Strategy API to allow for pluggable conflict resolution 3. Added TableService client API to schedule table services 4. Added Transaction Manager for wrapping actions within transactions	2021-03-16 16:43:53 -07:00
Raymond Xu	ab9933f206	[HUDI-1620] Add azure pipelines configs (#2582 )	2021-02-23 16:52:41 -08:00
n3nash	ffcfb58bac	[HUDI-1486] Remove inline inflight rollback in hoodie writer (#2359 ) 1. Refactor rollback and move cleaning failed commits logic into cleaner 2. Introduce hoodie heartbeat to ascertain failed commits 3. Fix test cases	2021-02-19 20:12:22 -08:00
pengzhiwei	37972071ff	[HUDI-1109] Support Spark Structured Streaming read from Hudi table (#2485 )	2021-02-17 03:36:29 -08:00
wangxianghu	7b2e658ac0	[MINOR] Add Jira URL and Mailing List (#2404 )	2021-01-27 19:48:42 -05:00
vinoth chandar	81836f0309	Removing spring repos from pom (#2481 ) - These are being deprecated - Causes build issues when .m2 does not have this cached already	2021-01-24 07:42:52 -08:00
Raymond Xu	84df26323d	[MINOR] Use skipTests flag for skip.hudi-spark2.unit.tests property (#2477 )	2021-01-24 21:36:41 +08:00
wangxianghu	d3ea0f957e	[HOTFIX] Revert upgrade flink verison to 1.12.0 (#2473 )	2021-01-22 10:55:46 -08:00
wenningd	976420c49a	[HUDI-1512] Fix spark 2 unit tests failure with Spark 3 (#2412 ) * [HUDI-1512] Fix spark 2 unit tests failure with Spark 3 * resolve comments Co-authored-by: Wenning Ding <wenningd@amazon.com>	2021-01-21 07:04:28 -08:00
Vinoth Chandar	3719e7b388	Moving to 0.8.0-SNAPSHOT on master branch.	2021-01-20 11:31:22 -08:00
Sivabalan Narayanan	b9c2856d16	[HUDI-1535] Fix 0.7.0 snapshot (#2456 ) * Revert "[MINOR] Bumping snapshot version to 0.7.0 (#2435)" This reverts commit `a43e191d6c`. * Fixing 0.7.0 snapshot bump	2021-01-19 12:20:43 -08:00
Sivabalan Narayanan	a43e191d6c	[MINOR] Bumping snapshot version to 0.7.0 (#2435 )	2021-01-16 09:56:28 -05:00
jshmchenxi	c3e9243ea1	[MINOR] Add maven profile to support skipping shade sources jars (#2358 ) Co-authored-by: Xi Chen <chenxi07@qiyi.com>	2021-01-03 23:19:48 -05:00
Danny Chan	76faf59652	[HUDI-1495] Upgrade Flink version to 1.12.0 (#2384 )	2020-12-29 10:15:43 +08:00
wenningd	286055ce34	[HUDI-1451] Support bulk insert v2 with Spark 3.0.0 (#2328 ) Co-authored-by: Wenning Ding <wenningd@amazon.com> - Added support for bulk insert v2 with datasource v2 api in Spark 3.0.0.	2020-12-25 09:43:34 -05:00
wenningd	fce1453fa6	[HUDI-1040] Make Hudi support Spark 3 (#2208 ) * Fix flaky MOR unit test * Update Spark APIs to make it be compatible with both spark2 & spark3 * Refactor bulk insert v2 part to make Hudi be able to compile with Spark3 * Add spark3 profile to handle fasterxml & spark version * Create hudi-spark-common module & refactor hudi-spark related modules Co-authored-by: Wenning Ding <wenningd@amazon.com>	2020-12-09 15:52:23 -08:00
wangxianghu	4d05680038	[HUDI-1327] Introduce base implemetation of hudi-flink-client (#2176 )	2020-11-18 17:57:11 +08:00
liujinhui	bfdce7b082	[HUDI-1193](Upgrade http dependency version) (#1970 )	2020-08-21 20:24:04 +08:00
Bhavani Sudha Saktheeswaran	4226d75144	Moving to 0.6.1-SNAPSHOT on master branch.	2020-08-14 12:54:15 -07:00
Udit Mehrotra	8d04268264	[HUDI-1174] Changes for bootstrapped tables to work with presto (#1944 ) The purpose of this pull request is to implement changes required on Hudi side to get Bootstrapped tables integrated with Presto. The testing was done against presto 0.232 and following changes were identified to make it work: Annotation UseRecordReaderFromInputFormat is required on HoodieParquetInputFormat as well, because the reading for bootstrapped tables needs to happen through record reader to be able to perform the merge. On presto side, this annotation is already handled. We need to internally maintain VIRTUAL_COLUMN_NAMES because presto's internal hive version hive-apache-1.2.2 has VirutalColumn as a class, versus the one we depend on in hudi which is an enum. Dependency changes in hudi-presto-bundle to avoid runtime exceptions.	2020-08-12 17:51:31 -07:00
Sivabalan Narayanan	9c24151929	[HUDI-1175] Commenting out testsuite tests from Integration tests until we investigate the CI flakiness (#1945 )	2020-08-10 21:00:57 -07:00
liujinhui	6b349b7711	[HUDI-210] Hudi Supports Prometheus Pushgateway (#1931 ) Co-authored-by: leesf <leesf@apache.org>	2020-08-09 15:29:54 +08:00

1 2 3 4 5

236 Commits