SimpleRDD Word Count
Problem statement
Use the semantics of a small RDD pipeline to count words across a batch of sentences. Split each sentence on one or more whitespace characters, ignore empty tokens, and treat words as case-sensitive.
Return one row [word, count] for each distinct word. Rows must follow the order in which each word first appears while scanning the sentences from left to right.
Function
simpleRddWordCount(sentences: String[]) → String[][]Examples
Example 1
sentences = ["hello world","hello spark","world hello"]return = [["hello","3"],["world","2"],["spark","1"]]Rows retain first-occurrence order.
Example 2
sentences = ["a a\tb","","b c"]return = [["a","2"],["b","2"],["c","1"]]Runs of spaces and tabs are delimiters and empty sentences add no words.
Example 3
sentences = ["Data data DATA"]return = [["Data","1"],["data","1"],["DATA","1"]]Word equality is case-sensitive.
Constraints
0 <= sentences.length <= 100000.- The total number of characters is at most
1000000. - Sentences contain printable ASCII characters and whitespace.