Monday, April 7, 2014

How to display SBT dependency graph ?

http://stackoverflow.com/questions/19606243/resolving-dependencies-in-creating-jar-through-sbt-assembly

Add to project/plugins.sbt:
addSbtPlugin("net.virtual-void" % "sbt-dependency-graph" % "0.7.4")
Add to build.sbt:
net.virtualvoid.sbt.graph.Plugin.graphSettings
Then run sbt dependency-graph.

Sunday, April 6, 2014

SparkR - Spark + R


# My test SparkR program - mySparkR.R

require(SparkR)

# this does not work since I don' have a cluster setup
# sc < - sparkR.init(master="spark://david-centos6:7077", sparkEnvir=list(spark.executor.memory="1g"))
sc < - sparkR.init(master="local[2]", sparkEnvir=list(spark.executor.memory="1g"))

lines < - textFile(sc, "hdfs://david-centos6:8020/user/david/data/result.txt")
       
words < - flatMap(lines,
   function(line) {
      strsplit(line, " ")[[1]]
})

wordCount < - lapply(words, function(word) { list(word, 1L) })

counts < - reduceByKey(wordCount, "+", 2L)
output < - collect(counts)

for (wordcount in output) {
  cat(wordcount[[1]], ": ", wordcount[[2]], "\n")
}
                                
# pur the input file into HDFS
> hadoop fs -put result.txt data

# Run SparkR to test it
> ./sparkR examples/mySparkR.R
- or To increase the memory used by the driver you can  -
> SPARK_MEM=1g ./sparkR examples/mySparkR.R


NOTE: SparkUI is at http://david-centos6:4040
reference: http://stackoverflow.com/questions/21677142/running-a-job-on-spark-0-9-0-throws-error

ERROR sometimes you will be experiencing:
"Initial job has not accepted any resources; check your cluster UI to ensure that workers are registered and have sufficient memory"

A simple test you can run to test if you have memory or other problems using Spark Shell
Try to run MASTER="local[2]" spark-shell on the same machine you're trying to run the code. And the same code in spark console: sc.parallelize(1 to 100).count

If the sufficient memory problem persist then you might want to try to add SPARK_WORKER_MEMORY=2g to the file tools/spark-0.9.0-incubating-bin-hadoop2/conf/spark-env.sh (Not sure if this help yet ???)



Friday, April 4, 2014

How to configure static IP on CentOS

https://gist.github.com/fernandoaleman/2172388
## Configure eth0 # # vi /etc/sysconfig/network-scripts/ifcfg-eth0 DEVICE="eth0" NM_CONTROLLED="yes" ONBOOT=yes HWADDR=A4:BA:DB:37:F1:04 TYPE=Ethernet BOOTPROTO=static NAME="System eth0" UUID=5fb06bd0-0bb0-7ffb-45f1-d6edd65f3e03 IPADDR=192.168.1.44 NETMASK=255.255.255.0 PREFIX=24 DNS1=208.67.222.222 DNS2=208.67.220.220 DNS3=208.67.222.220 HWADDR=00:24:7E:6D:C0:FB GATEWAY=192.168.1.1 ## Configure Default Gateway # # vi /etc/sysconfig/network NETWORKING=yes HOSTNAME=david-centos6 GATEWAY=192.168.1.1 ## Restart Network Interface # /etc/init.d/network restart ## Configure DNS Server # # vi /etc/resolv.conf nameserver 8.8.8.8 # Replace with your nameserver ip nameserver 192.168.1.1 # Replace with your nameserver ip
Note: don't forget to modify /etc/hosts by adding a new entry for 192.168.1.44 and remove 127.0.0.1 (because 127.0.0.1 will result in Spark Master refuse the connection request from Spark Worker)

Thursday, April 3, 2014

Sublime Text 3 for R


https://github.com/wuub/SublimeREPL
SublimeREPL
  1. Install Package Control. http://wbond.net/sublime_packages/package_control
  2. Install SublimeREPL
    1. Preferences | Package Control | Package Control: Install Package
    2. Choose SublimeREPL
  3. Restart SublimeText2
  4. Configure SublimeREPL (default settings in Preferences | Package Settings | SublimeREPL | Settings - Default should be modified in Preferences | Package Settings | SublimeREPL | Settings - User, this way they will survive package upgrades!

Enhanced-R package for Sublime Text 2/3

This package helps in writing R languages:
  • More comprehensive Indentation and Syntax
  • Send commands to different applications such as R GUI, Terminal and SublimeREPL.
  • Show function hint in status bar

You can search all the packages available for Sublime from: https://sublime.wbond.net/search/sublime


Wednesday, April 2, 2014

Embedding Scala in R vs embedding R in Scala

http://dahl.byu.edu/software/jvmr/dahl-payne-uppalapati-2013.pdf

1. Embedding Scala in R
[david@david-centos6 ~]$ R

Instantiating a Scala interpreter/compiler in R is accomplished as follows:
R> library("jvmr")
R> a < -  scalaInterpreter()

Multiple interpreters can be created and each maintains its own workspace and memory.
Scala code can be evaluated using the interpret function or its shorthand equivalent. The
following two lines of code are equivalent:
R> interpret(a,'val mu = 3')
R> a['val mu = 3']

Both the interpret function and its shorthand are capable of handling multi-line code:
R> a["val sigma = 2.5
val n = 10
"]

2. Embedding R in Scala
> JVMR_JAR=$(R --slave -e 'library("jvmr"); cat(.jvmr.jar)')
> scala -cp ".:$JVMR_JAR"

scala> import org.ddahl.jvmr.RInScala
import org.ddahl.jvmr.RInScala

scala> val R = RInScala()
R: org.ddahl.jvmr.RInScala = org.ddahl.jvmr.RInScala@40726e15

scala> R.eval("words <- his="" in="" made="" r="" span="" string="" was="">

scala> println(R.capture("words"))
[1] "This String was made in R"

3. Embedding Spark in R

Tuesday, April 1, 2014

Spark Stream

http://ampcamp.berkeley.edu/wp-content/uploads/2013/07/Spark-Streaming-AMPCamp-3.pptx
http://spark.incubator.apache.org/docs/latest/streaming-programming-guide.html

Unifying Batch and Stream Processing Models
- Spark program on Twitter log file using RDDs
val tweets = sc.hadoopFile("hdfs://...")
val hashTags = tweets.flatMap (status => getTags(status))
hashTags.saveAsHadoopFile("hdfs://...")
- Spark Streaming program on Twitter stream using DStreams
val tweets = ssc.twitterStream()
val hashTags = tweets.flatMap (status => getTags(status))

hashTags.saveAsHadoopFiles("hdfs://...")