Reading Sequence File in PySpark 2.0

Max picture Max · Jan 9, 2017 · Viewed 7.6k times · Source

I have a sequence file whose values look like

(string_value, json_value)

I don't care about the string value.

In Scala I can read the file by

val reader = sc.sequenceFile[String, String]("/path...")
val data = reader.map{case (x, y) => (y.toString)}
val jsondata = spark.read.json(data)

I am having a hard time converting this to PySpark. I have tried using

reader= sc.sequenceFile("/path","org.apache.hadoop.io.Text", "org.apache.hadoop.io.Text")
data = reader.map(lambda x,y: str(y))
jsondata = spark.read.json(data)

The errors are cryptic but I can provide them if that helps. My question is, is what is the right syntax for reading these sequence files in pySpark2?

I think I am not converting the array elements to strings correctly. I get similar errors if I do something simple like

m = sc.parallelize([(1, 2), (3, 4)])
m.map(lambda x,y: y.toString).collect()

or

m = sc.parallelize([(1, 2), (3, 4)])
m.map(lambda x,y: str(y)).collect()

Thanks!

Answer

zero323 picture zero323 · Jan 9, 2017

The fundamental problem with your code is the function you use. Function passed to map should take a single argument. Use either:

reader.map(lambda x: x[1])

or just:

reader.values()

As long as keyClass and valueClass match the data this should be all you need here and there should be no need for additional type conversions (this is handled internally by sequenceFile). Write in Scala:

Welcome to
      ____              __
     / __/__  ___ _____/ /__
    _\ \/ _ \/ _ `/ __/  '_/
   /___/ .__/\_,_/_/ /_/\_\   version 2.1.0
      /_/

Using Scala version 2.11.8 (OpenJDK 64-Bit Server VM, Java 1.8.0_111)
Type in expressions to have them evaluated.
Type :help for more information.
scala> :paste
// Entering paste mode (ctrl-D to finish)

sc
  .parallelize(Seq(
    ("foo", """{"foo": 1}"""), ("bar", """{"bar": 2}""")))
  .saveAsSequenceFile("example")

// Exiting paste mode, now interpreting.

Read in Python:

Welcome to
      ____              __
     / __/__  ___ _____/ /__
    _\ \/ _ \/ _ `/ __/  '_/
   /__ / .__/\_,_/_/ /_/\_\   version 2.1.0
      /_/

Using Python version 3.5.1 (default, Dec  7 2015 11:16:01)
SparkSession available as 'spark'.
In [1]: Text = "org.apache.hadoop.io.Text"

In [2]: (sc
   ...:     .sequenceFile("example", Text, Text)
   ...:     .values()  
   ...:     .first())
Out[2]: '{"bar": 2}'

Note:

Legacy Python versions support tuple parameter unpacking:

reader.map(lambda (_, v): v)

Don't use it for code that should be forward compatible.