lanyuanxiaoyao/hudi - hudi - Gitea: Git with a cup of tea

Author	SHA1	Message	Date
vinoth chandar	7a973a6944	[HUDI-159] Redesigning bundles for lighter-weight integrations - Documented principles applied for redesign at packaging/README.md - No longer depends on incl commons-codec, commons-io, commons-pool, commons-dbcp, commons-lang, commons-logging, avro-mapred - Introduce new FileIOUtils & added checkstyle rule for illegal import of above - Parquet, Avro dependencies moved to provided scope to enable being picked up from Hive/Spark/Presto instead - Pickup jackson jars for Hive sync tool from HIVE_HOME & unbundling jackson everywhere - Remove hive-jdbc standalone jar from being bundled in Spark/Hive/Utilities bundles - 6.5x reduced number of classes across bundles	2019-09-11 11:08:27 -07:00
leesf	5c2da6051e	[HUDI-225] Create Hudi Timeline Server Fat Jar	2019-08-29 20:03:06 -07:00
Balaji Varadarajan	5f9fa82f47	HUDI-124 : Exclude jdk.tools from hadoop-common and update Notice files (#858 )	2019-08-28 16:20:47 -07:00
vinoth chandar	cd090871a1	[HUDI-159]: Pom cleanup and removal of com.twitter.parquet - Redo all classes based on org.parquet only - remove unuused dependencies like parquet-hadoop, common-configuration2 - timeline-service does not build a fat jar anymore - Fix utilities and hadoop-mr bundles based on above	2019-08-25 16:01:14 -07:00
vinoth chandar	6edf0b9def	[HUDI-68] Pom cleanup & demo automation (#846 ) - [HUDI-172] Cleanup Maven POM/Classpath - Fix ordering of dependencies in poms, to enable better resolution - Idea is to place more specific ones at the top - And place dependencies which use them below them - [HUDI-68] : Automate demo steps on docker setup - Move hive queries from hive cli to beeline - Standardize on taking query input from text command files - Deltastreamer ingest, also does hive sync in a single step - Spark Incremental Query materialized as a derived Hive table using datasource - Fix flakiness in HDFS spin up and output comparison - Code cleanup around streamlining and loc reduction - Also fixed pom to not shade some hive classs in spark, to enable hive sync	2019-08-22 20:18:50 -07:00
Balaji Varadarajan	a4f9d7575f	HUDI-123 Rename code packages/constants to org.apache.hudi (#830 ) - Rename com.uber.hoodie to org.apache.hudi - Flag to pass com.uber.hoodie Input formats for hoodie-sync - Works with HUDI demo. - Also tested for backwards compatibility with datasets built by com.uber.hoodie packages - Migration guide : https://cwiki.apache.org/confluence/display/HUDI/Migration+Guide+From+com.uber.hoodie+to+org.apache.hudi	2019-08-11 17:48:17 -07:00
Balaji Varadarajan	ec965892b0	HUDI-149 - Remove platform dependencies and update NOTICE plugin	2019-08-05 08:57:15 -07:00
Luke Zhu	171901a9d0	Fix typo in hoodie-presto-bundle (#818 )	2019-08-01 08:51:57 -07:00
Balaji Varadarajan	6e0ff3a235	Generate Source Jars for bundle packages (#810 )	2019-07-30 18:17:14 -07:00
Balaji Varadarajan	a0d7ab2384	HUDI-70 : Making DeltaStreamer run in continuous mode with concurrent compaction	2019-06-18 17:48:14 -07:00
Balaji Varadarajan	479908fd20	HUDI-125 : Change License for all source files and update RAT configurations	2019-06-09 11:41:55 -07:00
Balaji Varadarajan	30b0f2636f	Changes related to Licensing work 1. Go through dependencies list one round to ensure compliance. Generated current NOTICE list in all submodules (other apache projects like flink does this). To be on conservative side regarding licensing, NOTICE.txt lists all dependencies including transitive. Pending Compliance questions reported in https://issues.apache.org/jira/browse/LEGAL-461 2. Automate generating NOTICE.txt files to allow future package compliance issues be identified early as part of code-review process. 3. Added NOTICE.txt and LICENSE.txt to all HUDI jars	2019-06-07 17:58:57 -07:00
guanjianhui	173e0b6be4	exlude fasterxml and parquet from presto bundle	2019-06-07 11:33:43 -07:00
Thinking	66893bfef2	fix spark-shell add jar problem jira link https://issues.apache.org/jira/browse/HUDI-101 issue link https://github.com/apache/incubator-hudi/issues/516#issue-386048519 when using spark-shell with hoodie save data like : ``` ./spark-shell --master yarn --jars /home/hdfs/software/spark/hoodie/hoodie-spark-bundle-0.4.8-SNAPSHOT.jar --conf spark.sql.hive.convertMetastoreParquet=false --packages com.databricks:spark-avro_2.11:4.0.0 ``` and ``` inputDF.write.format("com.uber.hoodie") .option("hoodie.insert.shuffle.parallelism", "1") // any hoodie client config can be passed like this .option("hoodie.upsert.shuffle.parallelism", "1") // full list in HoodieWriteConfig & its package .option(DataSourceWriteOptions.STORAGE_TYPE_OPT_KEY, HoodieTableType.COPY_ON_WRITE.name()) .option(DataSourceWriteOptions.OPERATION_OPT_KEY, DataSourceWriteOptions.UPSERT_OPERATION_OPT_VAL) // insert .option(DataSourceWriteOptions.RECORDKEY_FIELD_OPT_KEY, "_row_key") .option(DataSourceWriteOptions.PARTITIONPATH_FIELD_OPT_KEY, "partition") .option(DataSourceWriteOptions.PRECOMBINE_FIELD_OPT_KEY, "extend_deal_date") .option(HoodieWriteConfig.TABLE_NAME, "c_upload_code") .mode(SaveMode.Overwrite) .save("/tmp/test/hoodie") ``` It also report error `Invalid signature file digest for Manifest main attributes`. Need to scan all infected dependency.	2019-06-03 15:01:43 -07:00
Vinoth Chandar	7b4a28ecf8	Move depedency repos to https urls	2019-05-31 20:37:03 -07:00
Vinoth Chandar	acd74129cd	Create hoodie-utilities-bundle to host the shaded jar - hoodie-utilities can now be pulled in as compile time dependency - Lets users test their DeltaStreamer transformers for e.g - Tested the docker demo works & takes in the bundle - Doc changes to follow, to move DeltaStreamer commands to bundle jar	2019-05-30 22:46:24 -07:00
vinothchandar	66c0b81b49	[maven-release-plugin] prepare for next development iteration	2019-05-28 19:17:26 -07:00
vinothchandar	227785c022	[maven-release-plugin] prepare release hoodie-0.4.7	2019-05-28 19:17:15 -07:00
Balaji Varadarajan	64fec64097	Timeline Service with Incremental View Syncing support	2019-05-16 13:25:33 -07:00
vinothchandar	446f99aa0f	[maven-release-plugin] prepare for next development iteration	2019-05-14 07:29:22 -07:00
vinothchandar	cc38abecc8	[maven-release-plugin] prepare release hoodie-0.4.6	2019-05-14 07:29:11 -07:00
Abhishek Sharma	e2dcef8606	HUDI-101: added exclusion filters for signature files.	2019-05-07 18:35:18 -07:00
Omkar Joshi	738635306b	migrating kryo's dependency from twitter chill to plain kryo library	2019-05-06 20:32:00 -07:00
Balaji Varadarajan	36ef94004e	Fix Hive RT query failure in hoodie demo	2019-04-17 16:36:32 -07:00
Omkar Joshi	e35d24f31d	Revert "Replacing Apache commons-lang3 object serializer with Kryo serializer" This reverts commit `a6c45feb2c`.	2019-04-17 09:23:37 -07:00
Bhavani Sudha Saktheeswaran	83b6aa5e91	Fix multiple issues when using build_local_docker_images for setting up the demo Details here - https://issues.apache.org/jira/browse/HUDI-98	2019-04-15 10:10:05 -07:00
Balaji Varadarajan	b07110b9fd	Essential Hive packages missing in hoodie spark bundle	2019-04-09 21:42:42 -07:00
Omkar Joshi	a6c45feb2c	Replacing Apache commons-lang3 object serializer with Kryo serializer	2019-03-18 14:12:25 -07:00
Balaji Varadarajan	adc8cac743	Fix hive sync (libfb version mismatch) and deltastreamer issue (missing cmdline argument) in demo	2019-03-13 16:14:32 -07:00
vinothchandar	687395e40f	[maven-release-plugin] prepare for next development iteration	2019-02-27 07:16:27 -08:00
vinothchandar	bbf40ef987	[maven-release-plugin] prepare release hoodie-0.4.5	2019-02-27 07:16:15 -08:00
Bhavani Sudha Saktheeswaran	75c7a2622b	Create hoodie-presto bundle jar Exclude common dependencies that are available in Presto	2019-02-24 19:49:02 -08:00
Kent Yao	09f203d324	typo: bundle jar with unrecongnized variables	2019-02-13 16:46:11 +08:00
Balaji Varadarajan	3a0044216c	New Features in DeltaStreamer : (1) Apply transformation when using delta-streamer to ingest data. (2) Add Hudi Incremental Source for Delta Streamer (3) Allow delta-streamer config-property to be passed as command-line (4) Add Hive Integration to Delta-Streamer and address Review comments (5) Ensure MultiPartKeysValueExtractor handle hive style partition description (6) Reuse same spark session on both source and transformer (7) Support extracting partition fields from _hoodie_partition_path for HoodieIncrSource (8) Reuse Binary Avro coders (9) Add push down filter for Incremental source (10) Add Hoodie DeltaStreamer metrics to track total time taken	2019-02-11 18:22:05 -08:00
vinothchandar	7ba842c0fe	[maven-release-plugin] prepare for next development iteration	2018-09-28 11:27:00 +05:30
vinothchandar	5847b61f44	[maven-release-plugin] prepare release hoodie-0.4.4	2018-09-28 11:26:15 +05:30
Balaji Varadarajan	2728f96505	Add dummy classes to dump all classes loaded as part of packaging modules to ensure javadoc and sources jars are getting created	2018-09-18 09:24:33 +05:30
Vinoth Chandar	bd5af89f12	[maven-release-plugin] rollback the release of hoodie-0.4.4	2018-09-13 15:01:53 +05:30
Vinoth Chandar	d1cc864a43	[maven-release-plugin] prepare for next development iteration	2018-09-12 23:59:47 +05:30
Vinoth Chandar	b748bc836d	[maven-release-plugin] prepare release hoodie-0.4.4	2018-09-12 23:59:34 +05:30
Balaji Varadarajan	18a39715c9	Bump up versions in packaging modules and remove commons-lang3 dep	2018-09-11 11:03:30 +05:30
Vinoth Chandar	eca49a255e	Rebasing and fixing conflicts against master	2018-09-11 11:03:30 +05:30
Vinoth Chandar	a5359662be	Moving depedencies off cdh to apache + Hive2 support - Tests redone in the process - Main changes are to RealtimeRecordReader and how it treats maps/arrays - Make hive sync work with Hive 1/2 and CDH environments - Fixes to make corner cases for Hive queries - Spark Hive integration - Working version across Apache and CDH versions - Known Issue - https://github.com/uber/hudi/issues/439	2018-09-11 11:03:30 +05:30

1 2 3 4

193 Commits