Professional Documents
Culture Documents
• Databases
• MySQL, MS SQL Server,
PostgreSQL
• any JDBC (Oracle, DB2, …)
• Files
• CSV, txt
• Excel, Word, PDF
• SAS, SPSS
• XML, JSON
• PMML
• Images, texts, networks, chem
• Web, Cloud
• REST, Web services
• Twitter, Google
• Spark
• HDFS support
• Hive
• Impala
• Vertica
• In-database processing
• Preprocessing
• Row, column, matrix based
• Data blending
• Join, concatenate, append
• Aggregation
• Grouping, pivoting, binning
• Feature Creation and Selection
• Regression
• Linear, logistic
• Classification
• Decision tree, ensembles,
SVM, MLP, Naïve Bayes
• Clustering
• k-means, DBSCAN, hierarchical
• Validation
• Cross-validation, scoring, ROC
• Deep Learning
• Keras, DL4J
• External
• R, Python, Weka, H2O
• Interactive Visualizations
• JavaScript-based nodes
– Scatter Plot, Box Plot, Line Plot
– Networks, ROC Curve, Decision
Tree
– Adding more with each release!
• Misc
• Tag cloud, open street map,
molecules
• Script-based visualizations
• R, Python
• Database
• Files
• Excel, CSV, txt
• XML
• PMML
• to: local, KNIME Server, SSH-,
FTP-Server
• BIRT Reporting
KNIME
Explorer
Workflow Editor
Node
Recommendations
Node Description
Node Repository
Console
Outline
Recommendation engine
– Gives hints about which node use next in the workflow
– Based on KNIME communities' usage statistics
– Based on own KNIME workflows
The buttons in the toolbar can be used for the active workflow. The most
important buttons:
– Execute selected and executable nodes (F7)
– Execute all executable nodes
– Execute selected nodes and open first view
– Cancel all selected, running nodes (F9)
– Cancel all running nodes
Configured:
The node has been configured correctly, and can be executed.
Executed:
The node has been successfully executed. Results may be
viewed and used in downstream nodes.
Model
Flow Variable
Image
Database Database
Data Conection Query
• Right-click node
• Select Execute in context menu
• If execution is successful, status shows
green light
• If execution encounters errors, status
shows red light
• Right-click node
• Select Views in context menu
• Select output port to inspect execution results
Plot View
Data View
Node Execution Shift + F10 executes all configured nodes and opens all views
F9 cancels selected running nodes
Shift + F9 cancels all running nodes
Ctrl + Shift + Arrow moves the selected node in the arrow direction
Move Nodes and moves the selected annotation in the front or in the back
Ctrl + Shift +
Annotations of all overlapping annotations
PgUp/PgDown
F8 resets selected nodes
Ctrl + S saves the workflow
Workflow Operations
Ctrl + Shift + S saves all open workflows
Ctrl + Shift + W closes all open workflows
Meta-node Shift + F12 opens meta-node wizard
• Additional tasks:
– Sentiment analysis
– Visualization of documents
– Document clustering
Title
Rating
Author
Fulltext
Goal:
• Build a classifier to distinguish between reviews
about Italian or Chinese restaurants.
Review about
an Italian or a
Chinese
restaurant?
Classifi-
cation
• Document Cell
– Encapsulates a document
• Title, sentences, terms, words
• Authors, category, source
• Generic meta data (key, value pairs)
• Term Cell
– Encapsulates a term
• Words, tags
• Document table
– List of documents
• Bag of words
– Tuples of documents
and terms
• Document vectors
– Numerical
representations of
documents
• Open KNIME
• Import workflows from USB stick
Output port
Status
Node name
File path
Basic
Settings Advanced
Settings
Help
Preview button
File path
Sheet
specific
settings
Preview
Licensed under a Creative Commons Attribution- ®
55
Copyright © 2018 KNIME AG
6 Noncommercial-Share Alike license
https://creativecommons.org/licenses/by-nc-sa/4.0/
New Node: Table Reader
File path
Directory
Recursive
search
File
extensions
Meta data to
Extraction of extract
attachments
Title
Text
Category
Authors
Tokenizer
Import text
• Table Reader
• Row Filter
• Strings to
Documents
• Column Filter
Overwrite!
Set unmodifiable
in tagger nodes
Ignore
unmodifiability in
preprocessing
nodes
• Node Repository:
Other Data Types/Text Processing/Enrichment
• Available Tagger Nodes
– Stanford tagger
– Dictionary (& Wildcard) tagger
– OpenNLP tagger
– Abner tagger
–…
Number of
parallel
Model to
threads
use
Exact match
Dictionary or “contains”
column
Tag value to
Type of tag be assigned
to be assigned
Enrichment
• POS tagger
Enrichment
• File Reader
• Dictionary Tagger
Document
corpus StanfordNLP
NE model
Dictionary
Document
corpus
Dictionary
column
Tag type
and value
Use model
from input
port or built-
in models
• Trains NER model for latin and gallic names based on “De
Bello Gallico” from Julius Caesar.
• Node Repository:
Other Data Types/Text Processing/Preprocessing
• Available Preprocessing Nodes
– Stop Word Filter
– Snowball Stemmer
– Tag Filter
– Case Converter
– RegEx Filter
–…
Append
original
document
Ignore term
unmodifiability
Built-in stop
word lists
Custom stop
word list
Language
selection
Tag type
selection
Tag value
selection
Preprocessing
• Number Filter
• Punctuation Erasure
• Stop Word Filter
• Case Converter
• Snowball Stemmer
• POS Filter
• Node Repository:
Other Data Types/Text Processing/Transformation
• Available Transformation Nodes
– Bag of Words Creator
– Document Vector
– Strings to Document
– Sentence Extractor
– Document Data Extractor
–…
Documents
used to create
bag of words
Original
documents
can be
appended
Terms to
transform to
strings
Preprocessing II
• Bow Creator
• Term to String
• GroupBy
• Row Filter
• Reference Row Filter
Documents to
append to left
Create bit or of the created
numerical vector columns
vector
Reference
document vectors
Licensed under a Creative Commons Attribution- ®
112
Copyright © 2018 KNIME AG
13 Noncommercial-Share Alike license
https://creativecommons.org/licenses/by-nc-sa/4.0/
Document Vector Applier: Configuration
Use settings
from model
input
Include and
exclude lists of
features of the
reference
vectors
Dimensions of
document vectors
Hashing
function
Extracted field as
Document string column
column
Fields to
extract
• Node Repository:
Other Data Types/Text Processing/Frequencies
• Available Frequency Nodes
– TF
– IDF
– Ngram creator
–…
Appended column
with TF values
Appended column
with DF values
Appended column
with IDF values
Transformation
• TF
• Document Vector
• Document Data
Extractor
Methods:
• Decision Trees
• Neural Networks
• Naïve Bayes
• Logistic Regression
• Support Vector
Machine
• Tree Ensembles
Train
Training Model
Set Score
Apply
Model Model
Original
Data Set
Test
Set
Training and
Data Partitioning Model Evaluation
Applying Models
Train
Training Model
Set Score
Apply
Model Model
Original
Data Set
Test
Set
Test set
• C4.5 builds a tree from a set of training data using the concept
of information entropy.
• At each node of the tree, the attribute of the data with the
highest normalized information gain (difference in entropy) is
chosen to split the data.
• The C4.5 algorithm then recourses on the smaller sub lists.
Train
Training Model
Set Score
Apply
Model Model
Original
Data Set
Test
Set
False Positives
False Negatives
True Negatives
From the
confusion matrix
Classification
• Color Manager
• Column Filter
• Partitioning
• Decision Tree Learner
• Decision Tree
Predictor
• Scorer
Classification II
Methods:
• Predictive modeling
• Dictionary based
• Deep parsing
• …
Predictive modeling:
• Build classifier to distinguish between positive and negative
reviews.
– “Ah, Moonwalker, I'm a huge Michael Jackson fan, I grew up with his
music, Thriller was actually the first music video I ever saw apparently.
…”
– “This film has a very simple but somehow very bad plot. …”
Positive or
negative?
Classification
• Strings to document
• Preprocessing nodes
• Bag of words creation,
grouping, counting,
and filtering
• Vector creation
• Model training and
scoring
Dictionary based:
• Use a custom dictionary to count positive and
negative words.
• Compute sentiment score to predict sentiment
label.
Classification
• Strings to
Documents
• Dictionary Tagger
• Bag of words, TF,
and GroupBy for
counting
• Pivoting
• Math Formula
• Rule Engine
• Scorer
• Node Repository:
Other Data Types/Text Processing/Misc
• Available Visualization Nodes
– Document Viewer
– Tag Cloud
Term column
and frequency
column
Scaling of font
size: linear,
log, exp
Visualization
• Decision Tree
Learner
• Tag Cloud
• (Optional Coloring)
– TF, Document Data
Extractor, Group By,
Pivoting, Math
Formula, Color
Manager
Document
column
Tagged terms
can be hilited
Author, category,
meta information,
…
Licensed under a Creative Commons Attribution- ®
167
Copyright © 2018 KNIME AG
11 Noncommercial-Share Alike license
https://creativecommons.org/licenses/by-nc-sa/4.0/
Section Exercise
Visualization
• Document Viewer
• Supplementary Workflows/
– R Theme River (R plot)
– Twitter Word Tree (JavaScript view)
Methods:
• Hierarchical
clustering
• K-Means / Medoids
• Density based
• …
Name of Distance
distance measure
column
Columns to use
for distance
computation
Distance
function
(optional) 178
Licensed under a Creative Commons Attribution- ®
Copyright © 2018 KNIME AG
8 Noncommercial-Share Alike license
https://creativecommons.org/licenses/by-nc-sa/4.0/
Hierarchical Clustering (DistMatrix): Configuration
Distance
column Linkage
type
Threshold or
cluster count
based assignment
Hierarchy of
data points Illustration of
dendrogram
Data e.g.:
document
vectors Assignment of
clusters
Licensed under a Creative Commons Attribution- ®
183
Copyright © 2018 KNIME AG
13 Noncommercial-Share Alike license
https://creativecommons.org/licenses/by-nc-sa/4.0/
New Node: k-Medoids
• Similar nodes:
– k-Means
– Fuzzy c-Means
Distance matrix
column Cluster count k
Random seed
for reproducible
results
Clustering
• Distance Matrix
Calculate
• Hierarchical
Clustering
– Cluster View
– Cluster Assigner
• k-Medoids
Try Catch
Block
Sentiment analysis of
users
Licensed under a Creative Commons Attribution- ®
198
Copyright © 2018 KNIME AG
10 Noncommercial-Share Alike license
https://creativecommons.org/licenses/by-nc-sa/4.0/
Social Media Analysis
• Interaction network of
characters.
• Border color indicates
family assignment
• Node size is related to TF
of character names