Search in book...
Toggle Font Controls
Create new playlist

Name your new playlist

Playlist description (optional)
Sign In

Email address

Password

Forgot Password?

or

Continue with Facebook

Continue with Google
Sign Up

Full Name

Email address

Confirm Email Address

Password

or

Continue with Facebook

Continue with Google

Summary

In this chapter, we started with a Hello World program and setting up an IDE (Eclipse) for executing Spark Jobs. Then we discussed various RDD Transformation and various common transformations such as map, flatMap, mapToPair, and so on. We also explored commonly used RDD actions and some of the use cases associated with it. We also gained some understanding in improving Spark Job performance by using Apache Spark's inbuilt cache and persist mechanism.

The next chapter will focus on the interactional aspect of Apache Spark as far as the Data and Storage layer is concerned. We will learn about Spark integration with external storage systems such as HDFS, S3 etc and its ability to process various data formats such as xml, json etc.

..................Content has been hidden....................

You can't read the all page of ebook, please click here login for view all page.

3.17.76.72

Table of Contents for Summary

Create new playlist

Sign In

Sign Up

Table of Contents for
Summary