37

edf.select("x").distinct.show() shows the distinct values that are present in x column of edf DataFrame.

Is there an efficient method to also show the number of times these distinct values occur in the data frame? (count for each distinct value)

Leothorn
  • 1,295
  • 1
  • 22
  • 43

6 Answers6

71

countDistinct is probably the first choice:

import org.apache.spark.sql.functions.countDistinct

df.agg(countDistinct("some_column"))

If speed is more important than the accuracy you may consider approx_count_distinct (approxCountDistinct in Spark 1.x):

import org.apache.spark.sql.functions.approx_count_distinct

df.agg(approx_count_distinct("some_column"))

To get values and counts:

df.groupBy("some_column").count()

In SQL (spark-sql):

SELECT COUNT(DISTINCT some_column) FROM df

and

SELECT approx_count_distinct(some_column) FROM df
Alper t. Turker
  • 32,514
  • 8
  • 78
  • 112
zero323
  • 305,283
  • 89
  • 921
  • 912
14

Roughly speaking, how it works:

enter image description here

enter image description here

Saurav Sahu
  • 11,445
  • 5
  • 53
  • 73
10

Another option without resorting to sql functions

df.groupBy('your_column_name').count().show()

show will print the different values and their occurrences. The result without show will be a dataframe.

Antoni
  • 2,359
  • 18
  • 21
6
import org.apache.spark.sql.functions.countDistinct

df.groupBy("a").agg(countDistinct("s")).collect()
Community
  • 1
  • 1
user10232195
  • 61
  • 1
  • 3
3

If you are using Java, then import org.apache.spark.sql.functions.countDistinct; will give an error : The import org.apache.spark.sql.functions.countDistinct cannot be resolved

To use the countDistinct in java, use the below format:

import org.apache.spark.sql.functions.*;
import org.apache.spark.sql.*;
import org.apache.spark.sql.types.*;

df.agg(functions.countDistinct("some_column"));
ForeverLearner
  • 1,593
  • 2
  • 24
  • 47
1
df.select("some_column").distinct.count
Petter Friberg
  • 20,644
  • 9
  • 57
  • 104
shengshan zhang
  • 498
  • 8
  • 16