Direkt zum Hauptbereich

Posts

#LinearRegression has no existence - #Hegel would say

In #datamining, we are bored by linear regression. It does not work very well and I have personally never seen a qqplot that was strictly on a line, anyway. But besides this practical approach, linear regression (still very popular in social science) seems to have a strong theoretical foundation in the central limit theorem. But here comes an argument derived from Hegel's Logic, why linear regression has no existence at all: Hegel distinguishes between "bad" infinity and real infinity. Bad infinity is the result of infinite progress. Take 2:7 as an example. If you want to solve this division, you will end up with 0.285714285714... and so on (infinite). Or think about the rock, paper, scissor game: You want to take the good old rock, but then you think, your opponent might think that you want to take rock and therefore will choose paper. So you want to choose scissor, but then you think, your opponent knows that you wanted to take rock but he realized that you had rea...

How to select specific random rows from #MySQL within #R

When you work with big data sets it is a good idea to store everything in a MySQL database. Doing statistics, you often will need a random sample of your data, perhapse from a subpopulation of the data. The normal way is to run an SQL-querry with the ORDER BY RAND() addition. Unfortunately, this is slow as hell. A better way is to load all autoids or rownumbers of the subpopulation first, select a random sample in R and then run a second request where the sample is directly addressed. This is very fast even with huge data sets and has the advantage that you can control the randomness via set.seed. Here is an example: require ( RMySQL ) # set up RMySQL-connection con <- dbConnect ( RMySQL ::MySQL ( ) , host = "myhost" , user = "myusername" , password = "mypassword" , dbname= "mydb" )   # select the auto-ID or Rownumber where the column "col" has the value "select" IDs <- dbSendQuery ( con ,...

#I-Hegel: Why there will never be artificial inteligence (#AI)

... but perhaps something more powerfull. Lately, I have been thinking about AI a lot. Right now, I am readig Hegel again, and I am trying to do it seriously (sorry guys, I do not know if this could be done in English...). What strikes me is this: There is no intelligence. Intelligence is a force ("Kraft") that is said to lead to different expressions ("Äußerung"). Think about a guy who plays chess on a grandmaster's level and is able to solve any mathematical equation with ease. He is very intelligent, isn't he? The problem is that we do not know anything about the force exept for the expressions. We can not find intelligence in anyone without him expressing something we declare to be intelligent. Therefore, the expression and the force cannot be differentiated in reality. Someone does something we label intelligent and therefore we say that the mystique force of intelligence is somewhere in him or her. Hegel tells us that this is a wrong judgement: T...

Handling #bigdata in R (2): Cuda and rpud

Taken into acount, how long it took to just load the data from the kaggle competition , working with it is kinda scarry. Therefore, I gonna speed up my system by using GPU computing in R. Modern graphic cards are very powerful and can - in principle - work like a highperformance cluster. Following this brilliant tutorial by Chi Yau, I managed to get Cuda running on my Ubuntu machine. It took some time, especially the "make" command at the end. But, as you will see shortly, it is absolutely worth the affort. The second part of the tutorial shows how to install rpud (don't try to install it directly from R-Studio, but follow these steps.) Again, it took some time, especially until I realized that the type="source" parameter has to be added when installing the packadge. Finaly, everything is working and I followed Chi Yau's example and calculated a distant matrix for datasets with huge number of vectors. The gain in speed is unbelivable! And here is...

Handling #bigdata in R (1): data.table

How big is big? Are you fit for real big data? These are some questions I am thinking about. Luckily, there is a kaggle competition going on, with the aim to predict the click through rate in a huge dataset of webpage visits. The task is to predict the probability that 4.5 Million users are clicking on an advertisment. The training dataset contains of 40 Million (!!!) lines of user data. In this little series I will share my experience in trying to handle this mess. First problem: How to load the data? Loading the csv file in a normal way takes much too long. The package "data.table" includes the fread function, which is much faster. By setting colClasses to "character" all columns are loaded as character class library ( data.table ) train <- fread ( "train.csv" , colClasses= "character" )   Reading nearly 6 GB will take some time, nevertheless...

#Phreaking is back - on android!

Phreaking or phone freaking has been the begining of the hacking culture. "The term first referred to groups who had reverse engineered the system of tones used to route long-distance calls. By re-creating these tones, phreaks could switch calls from the phone handset, allowing free calls to be made around the world." ( Wikipedia ) OK, you cannot phone for free by whisteling in your android device. But try this: Type *#*#4636#*#* in your smartphone and you find a secret menu. There are many other secret codes, including functions to reset the entire device in fabrique state. Even more spooky: Some webpages simulate the dialing of these numbers. So surfing the web with your smartphone may cause seriouse trouble.   Taken from Wizzywig .

Using ngrams with #RTextTools

This littel example shows a workaround for a bug in RTextTools. Using ngramLength would lead to an error. But we can use the RWeka library and tm, then go back to RTextTools: library ( RTextTools ) texts <- c ( "This is the first document." ,   "Is this a text?" , "This is the second file." ,   "This is the third text." ,   "File is not this." ) library ( RWeka ) library ( tm ) TrigramTokenizer <- function ( x ) NGramTokenizer ( x , Weka_control ( min = 3 , max = 3 ) ) dtm <- DocumentTermMatrix ( Corpus ( VectorSource ( texts ) ) , control = list ( weighting=weightTf , tokenize = TrigramTokenizer ) ) as.matrix ( dtm ) isText <- c ( T , F , T , T , F ) container <- create_container ( dtm , isText , virgin=F , trainSize= 1 : 3 , testSize= 4 : 5 )   models=train_models ( container , algorithm= c ( "SVM" , ...