The new sparklyr package is a native dplyr interface to Spark, according to RStudio. After installing the package, users can “interactively manipulate Spark data using both dplyr and SQL (via DBI), ...