Philippe ROSSIGNOL
@enahwe
About
I'm a Big Data consultant mainly motivated around Spark and Hadoop technologies
Skills & Technologies
Projects & Repositories
Recent public projects and repositories from this profile.
Csv2Hive
ShellCsv2Hive is an useful CSV schema finder for the Big Data. It discovers automatically schemas in big CSV files, generates the 'CREATE TABLE' statements and creates Hive tables. You don't need to writes any schemas at all. Csv2Hive is a really fast solution for integrating the whole CSV files into your DataLake.
KafkaGust
JavaKafkaGust is a flexible and useful tool to quickly test high data volumes with Apache Kafka. Apache Kafka that is used in the world of Big Data (e.g : with Storm, Cassandra, Hadoop, etc.), is a fast and scalable publish-subscribe messaging that can handle durably hundreds of megabytes of reads and writes per second from thousands of clients.
Spark-DIL-FTP
A Spark Application for FTP Data Ingestion
Spark-DIL
PythonA Spark Library for Data Integration