Showing posts with label column store. Show all posts
Showing posts with label column store. Show all posts

Wednesday, January 2, 2013

BigData Analysis with Project Spark and Shark


SPARK :
  • Developed by AMPLabs, UC Berkeley
  •  Developers Michael Franklin and Matei Zaharia 
  •  Alternative to MapReduce parallel processing engine.
  •  In-memory storage for very fast iterative queries removing temporary writes of intermediate data like MapReduce jobs. 
  •  After each map and shuffle the data is written to local disk in Hadoop. Which increases the further  execution time. This bottle neck removed in SPARK by making the results available in the memory itself.
  • Spark writes data to RDD (Resilient Distributed Datasets) which can live memory and hence Spark provides the necessary execution improvements.
  • Up to 100x faster than Hadoop.
  • Compatible with existing Hadoop ecosystem and works well with existing HDFS systems. 
  • Spark can co-exist with existing Hadoop cluster using Mesos cluster manager. 
  • It is better suited for iterative algorithms like Logistic Regression and Matrix Factorization compared plain data processing algorithms.
  • Developed by Scala and provides clean APIs in Java and Scala. Python APIs will be added soon.
SHARK :
  • Meant for Hive replacement with high degree of speed improvement.
  • Built on top of SPARK data-parallel execution engine.
  • Uses SQL like declarative language and works on SPARK infrastructure.
  • Can execute complex queries using JOINs and GROPU BY
  • Uses column-oriented store to improve performance. The columnar compression provides better reduction in storage.
  • All the queries run in memory to improve the performance.
  • Shark provides descent integration with Machine Learning using (Resilient Distributed Datasets). User can call these functions using SQL like syntax. This minimizes the complexity involved in using machine language.
  • The entire software stack ( SHARK + SPARK + ecosystem ) is called as BDAS ( Berkeley Data Analysis Stack ).


Monday, December 10, 2012

Trevni : A columnar file format for Cloudera Impala


Trevni is columnar file format developed by Doug Cutting for storing data in columnar format and will be core storage engine and part of  Cloudera Impala project. Project impala delivers real time queries on Hadoop file system.

Trevni features :
  1. Inspired by CIF/COF based column oriented database architecture.  CIF based columnar architecture works well with MapReduce.
  2. Stores data based on columns which provides good compression of data as data stored in single column will have same kind of data. Retrieval of data will be fast as the minimal scanning required for accessing the data within the same column as compared row store.
  3. To achieve scalable, distributed query evaluation, data sets are partitioned into row groups containing distinct collection of rows. Then each row group stores data vertically like column store. To understand more see below figure 1.
  4. Maximizes the size of row groups in order to reduce the Disc IO seeking latency.  Each row group size can be > 100 mb. This will help in reading sequentially to reduce the disk IO.
  5. Each row group will be written as separate file. All values of a column will be written in contiguously to get optimized IO performance.
  6. Reducing no of row groups results in reducing the no of HDFS file created and hence it reduces the load on the name node. So it is better to have few files per data set means fewer row groups.
  7. Allows dynamic access of data within row group. It also supports co-location of columns with in row group as per CIF storage architecture.   
  8. It also supports nested column structure for semi structured data in the form of arrays and dictionaries.
  9. Application specific data will be maintained at every level like file, column and block. Check sums have been used at block level for providing data integrity.
  10.  Provides many data type support like int, long , float , double , string and byte data type for complex aggregated data. It also supports NULLs and NULL occupies zero bytes which is one of the key differences between column storage and row storage to save disk space. 

Figure - 1 : Illustrates the row group concept in columnar store.