Direkt zum Hauptbereich

Posts

#k-means #clustering

Hier kommt das versprochene Beispiel zu k-means Clustering. Angelehnt ist es an Ledolter, Johannes. 2013. Business analytics and data mining with R. Hoboken, NewJersey: Wiley . df = read.csv ( "http://www.biz.uiowa.edu/faculty/jledolter/DataMining/protein.csv" ) View ( df ) #what is the right number of clusters? #One idea is to look at the sum squared error (SSE) for #each possible number of clusters.   #Formula to calculate SSE: wss <- ( nrow ( df ) - 1 ) * sum ( apply ( df [ , - 1 ] , 2 , var ) )   for ( i in 2 : 20 ) wss [ i ] <- sum ( kmeans ( df [ , - 1 ] , centers=i ) $withinss ) plot ( 1 : 20 , wss , type= "b" , xlab= "Number of Clusters" , ylab= "Within groups sum of squares" )     #The "elbow" now indicates the optimal number of clusters. The #idea is that at a certain point additional clusters do not #reduce the SSE much... #In this case, we might decide to take 5 clusters.   #But ...

#NRW - Tweets aus Nordrhein-Westfalen in #R saugen

In diesem Post zeige ich ein kurzes Beispiel, wie man das streamR-Package nutzen kann, um Tweets aus NRW zu speichern. Voraussetzung ist, dass man ein consumer secret und einen consumerkey von Twitter hat. Das Beispiel aus der streamR-Doku ist im Prinzip richtig. Nur muss man inzwischen https anstelle von http verwenden. Hier also "stuff that works", wie Guy Clark sagen würde. Wichtig ist, dass man eine gültige cacert.pem Datei im Arbeitsverzeichnis hat. Eine entsprechende Datei findet sich hier .  #load libraries library(streamR) library(ROAuth) library(twitteR) requestURL <- "https://api.twitter.com/oauth/request_token" accessURL <- "https://api.twitter.com/oauth/access_token" authURL <- "https://api.twitter.com/oauth/authorize" #own consumerKey and Secret is needed! consumerKey <- "xxxxxyyyyyzzzzzz" consumerSecret <- "xxxxxxyyyyyzzzzzzz111111222222" Cred <- OAuthFactory$new(consumerKey=consumerKe...

Love your #nearest neighbor

In this post the k-nearest neighbors algorithm is used to classify Twitter-data. We have 1000 tweets about fracking which are labeld "c" = contra fracking, "p" = pro fracking, "n" = neutral to fracking, "nr" = not related to fracking. The data can be found here . The data is already cleaned: Missing values have been coded as -9999, than all numeric variables have been normalized with this R-function: normalize = function(x) (x-min(x))/(max(x)-min(x)) Dates are transformed with, as can be seen here . The k-nearest-neighbors algorithm is a "lazzy" learner. It does not make any assumtpions about any distributions (non-parametric approach!). It simply calculates the distance in the feature space. Therefore, it does not "learn" anything because it won't come up with any abstraction. We want to test, if k-nearest-neighbors can identify the tweets just by the meta-data (without looking at the text...). Here is the R-scr...

#Datamining-Seminar Projekte

Liebe Studierende, bitte überlegt euch für Montag ein Projekt, das ihr gerne mit Data-Mining bearbeiten wollt. Außerdem solltet ihr euch schon einmal mit möglichen Daten vertraut machen. Hier ein paar Anregungen, wo Sozialwissenschaftler Datensätze herbekommen können: http://www.policyagendas.org/ Hervorragende Datensammlung zu US-Budgets und politischen Aufmerksamkeitsvariablen. Ähnliche Daten für andere Länder gibt es hier: http://www.comparativeagendas.info/ Daten der Industrienationen hat die OECD: http://stats.oecd.org/ Europadaten gibt's bei Eurostat: http://epp.eurostat.ec.europa.eu/portal/page/portal/statistics/search_database Deutschlanddaten bei destatis: https://www-genesis.destatis.de/genesis/online Regionale Daten aus Deutschland: https://www.regionalstatistik.de/genesis/online/logon Die US-Regierung hat eine open-data-Initiative: http://data.gov/ und Deutschland auch (allerdings mit viel weniger Daten): https://govdata.de/ Daten zur ameri...

Group Games Predictions

And the graph for the group games:

New Graphs Soccer Prediction

Here are some updated graphs of the random forest prediction: The lines show which teams will win more than 50 per cent (33 per cent) of their matches. Of course, this means a lot of randomness. As we have seen, even Spain does not win every match... and if you loose the wrong match...

Rank-Order of Soccer World Cup Predictions

In this post I start to present some simple graphs to visualize the results from the random forest prediction. Here, we see the mean of the winning probability of all teams in increasing order: