Overview

tfidf-cascalog

Implements a portion of the TF-IDF algorithm. Takes Avro based file as input, calculates TF, DF and D portions in a batch mode and places the results into a cassandra table. This will later be read by a storm-trident DRPC query and combined with the realtime data to form a complete view of the world.

Usage

Ensure that both hadoop and cassandra are started, then:

lein deps
lein compile
lein uberjar

copydata.sh
hadoop jar ./target/tfidf-cascalog-0.1.0-SNAPSHOT-standalone.jar data/document.avro data/en.stop 127.0.0.1

Obviously replacing the IP address with the appropriate cassandra IP address.

Tip: Filter by directory path e.g. /media app.js to search for public/media/app.js.
Tip: Use camelCasing e.g. ProjME to search for ProjectModifiedEvent.java.
Tip: Filter by extension type e.g. /repo .js to search for all .js files in the /repo directory.
Tip: Separate your search with spaces e.g. /ssh pom.xml to search for src/ssh/pom.xml.
Tip: Use ↑ and ↓ arrow keys to navigate and return to view the file.
Tip: You can also navigate files with Ctrl+j (next) and Ctrl+k (previous) and view the file with Ctrl+o.
Tip: You can also navigate files with Alt+j (next) and Alt+k (previous) and view the file with Alt+o.