Hadoop, Spark and Kafka have already had a defining influence on the world of big data, and now there’s yet another Apache project with the potential to shape the landscape even further: Apache Arrow.

The Apache Software Foundation on Wednesday launched Arrow as a top-level project designed to provide a high-performance data layer for columnar in-memory analytics across disparate systems.

Based on code from the related Apache Drill project, Apache Arrow can bring benefits including performance improvements of more than 100x on analytical workloads, the foundation said. In general, it enables multi-system workloads by eliminating cross-system communication overhead.

To read this article in full or to leave a comment, please click here

Source: COMPUTER WORLD