Showing posts with label data science. Show all posts
Showing posts with label data science. Show all posts

Wednesday, May 10, 2017

Not enough RAM for data frame? Maybe Dask will help?

What environment do we like?


When we are doing data analysis we would like to do it in as much interactive manner as possible. We would like to get results as soon as possible and test ideas without unnecessary waiting. We like to work in that manner because great proportion of our work leads to dead ends. As soon as we get there we can start thinking about other approach to problems. When you were working with R, Python or Julia backed by Jupyter Notebook, you are familiar with this workflow. And everything works well when your workstation has enough RAM to handle all data, its processing, cleaning, aggregation and transformations. Even Python, which is considered as rather slow language performs nicely here. Especially when you are using strongly optimized modules like NumPy. 

But what with case, where data which we want to process is larger than available RAM? Well, there are three simple outcomes of your attempts: 1) your operating system will start to swap and possibly will be able to complete ordered task, but it will take way more time than expected; 2) you will receive warning or error message from functions which can estimate needed resources - rare case; 3) your script will use all available RAM and swap and will hang your operating system. But are there maybe any other options?

What can we do?


According to this (Rob Story: Python Data Bikeshed) and this (Peader Coyle - The PyData map: Presenting a map of the landscape of PyData tools) presentations, there are many possibilities to build tailored data processing solution. From all of those modules, Dask looks especially interesting to me, because it mimics Pandas naming convention and functions behavior.

Dask offers three data structures - array, data frame and bag. Array structure is designed to implement methods known from NumPy array. So everyone experienced with NumPy will be able to use it without much problems. The same approach was used to implement data frame which is inspired by Pandas data frame structure. Bag on the other hand is designed to be equivalent of json dictionaries or other Python data structures.

What is key difference between Dask data structures and their archetypes? Dask structures are divided into small chunks and every operation on them is evaluated when its needed. It means that when we have data frame and series of transformations, aggregations, searches and similar operations Dask will calculate what and when to do with each chunk and will take care of executions of those operations and do garbage collection immediately. If we would like to do that in original Pandas data frame it will have to fit into RAM entirely, and every of steps in pipeline will also have to store its results in RAM even if some operations could be executed with inplace=True parameter.

How can we use it? As I mentioned, Dask data structures were designed to be "compatible" with NumPy and Pandas data structures. So, if we check data frame API reference we will see that many Pandas methods was re-implemented with the same arguments and results.

But the problem is that not all original methods from NumPy and Pandas are implemented in Dask. So it is not possible to blindly substitute Pandas with Dask and expect that code will work. On the other hand, in cases where you are unable to read your data at all, it might be worth to spend some time to rework your flow and adjust it to Dask.
 
Second problem with Dask it that despite it tries to execute various operations on chunks in parallel, it may take more time to produce final results than in simple all in RAM data frame. But if you know time execution characteristic of your scripts you can try to substitute some heavy time parts of it and compare with Dask execution. Maybe it will make sense.

The future?


What is the future of such modules? It depends. One my ask "Why to bother when we have PySpark?". And this is valid question. I would say, that Dask and similar solutions fits nicely in niche where data is to big for RAM but on the other hand it fits on directly attached hard drives. If data fits nicely in RAM I wouldn't bother to work with Dask and similar modules - I would just stick to plain old and good NumPy and Pandas. Also if I had to deal with such amount of data that wouldn't fit on available attached to workstation disks, I would consider going into big data solutions which also gives me redundancy over hardware failures. But Dask is still very cool, and at least worth to test.

I hope that Dask developers will port more of NumPy and Pandas methods into. I also saw some works towards integrating it with Scikit-Learn, XGBoost and TensorFlow. You should definitely check this out when you will be considering buying more RAM next time ;).

Friday, March 31, 2017

Bike Sharing Demand #01


Bike sharing system is one of most cool feature of modern city. Such system allows citizens and tourists to easily rent bikes, and return them in different places that they were rented. In Wrocław, it is even free to ride for first 20 minutes, so if your route is so short, or you can spot automated bike stations in this time interval, you can use rented bikes practically for free. It is very tempting alternative to crowded public transportation or private cars on jammed roads.

On 28-th May 2014 Kaggle started knowledge competition, in which goal was to predict bike rental number in city bike rental system. Bike system is owned by Capital Bikeshare which describe themselves: "Capital Bikeshare puts over 3500 bicycles at your fingertips. You can choose any of the over 400 stations across Washington, D.C., Arlington, Alexandria and Fairfax, VA and Montgomery County, MD and return it to any station near your destination."

Problem with bike sharing system is that, it need to be filled with ready to borrow bikes. Owner of such system need to estimate demand for bikes and prepare appropriate supply for them. If there will be not enough bikes, system will generate more disappointment and won't be popular. If there will be to much unused bikes, it will generate unnecessary maintenance costs on top of initial investment cost. So it seems that finding good estimate for renting demand could improve customer satisfaction and reduce unnecessary spendings.

How to approach such problem? Usually, first step should be dedicated to get some initial and general knowledge about available data. It is called Exploratory Data Analysis. I will perform EDA on data available in this competition.

As we can see, in train data we have three additional columns: {'casual', 'count', 'registered'}. Our goal its to predict 'count' value for each hour for missing days. We know that 'casual' and 'registered' should sum nicely to total 'count'. We can also observe their relations on scatter plots. 'registered' values seems to be nicely correlated with 'count', but 'casual' are also somewhat related. This plot can give idea, that instead of calculating 'count' one may calculate 'registered' and 'casual' and based on this numbers submit total 'count'.

Every data point is indexed by round hour in datetime format, so after splitting it to date and time components we have sixteen columns. We can easily generate histogram for each of them. By visually examining this histogram we can point to some potentially interesting features: '(a)temp', 'humidity', 'weather', 'windspeed', and 'workingday'. Are they important? We don't know that now.

We can examine more those features and pick those which have rather discrete values which unique count is less or equal 10 (my arbitrary guess). I will sum them for each hour for each unique value. Sounds complicated? Maybe plots will bring some clarity ;). First plot shows aggregation with keeping 'season' information. We have clear information that at some hours there were twice times more bike borrowing for season '3' than for season '1'. It is not so surprising if we assume that season '3' is summer. It will automatically lead to '1' being winter, and it is not true according to data description: season -  1 = spring, 2 = summer, 3 = fall, 4 = winter. Fascinating.

Second feature which might be interesting in this analysis is 'weather'. This feature is described as: weather - 1: Clear, Few clouds, Partly cloudy, Partly cloudy; 2: Mist + Cloudy, Mist + Broken clouds, Mist + Few clouds, Mist; 3: Light Snow, Light Rain + Thunderstorm + Scattered clouds, Light Rain + Scattered clouds; 4: Heavy Rain + Ice Pallets + Thunderstorm + Mist, Snow + Fog. It seems that 'weather' value corresponds with not niceness of weather conditions. Higher value - worst weather. We can see that on histogram and on aggregated hourly plot.
 
Another interesting information might be extracted from calculated feature 'dayofweek'. For each date there was day of week calculation and its results were encoded as 0 = Monday up to 6 = Sunday. As we can see on related image, there are two different trends on bike sharing based on day of week. There are two days which build first trend. They are '5' and '6' which means 'Saturday' and 'Sunday'. In western countries those days are part of weekend and are considered as non working days for most of population. Rest of days, which are all building second trend, are considered as working days. We can easily spot peaks on working days, which are by my guess related to traveling to and from work place. In second "weekend" trend we can observe smooth curves which are probably reflecting general humans activity over weekend days.
OK, it seems to be good time to examine correlations in this data set. Lets start with numerical and categorical correlations against our target feature 'count'. It is not surprising that 'registered' and 'casual' features are nicely correlated, we saw that earlier. 'atemp' and 'temp' seems to also correlate at some level. Rest of features have rather low correlation, but it seems, that there is 'humidity' among them which might also be worth of consideration for further possible investigations.

What are overall every to every feature correlations? We can examine it visually on correlations heatmap. On the plot, we can spot correlations between 'atemp' and 'temp'. In fact 'atemp' is defined as "feels like" temperature in Celsius. So it is subjective, but it correlates nicely with objective temperature in Celsius. Other correlation is in "month" and "season" which is not surprising at all. Similar situation we can observe with 'workingday' and 'dayofweek'.

So how much not redundant information do we have in this data set? After removing target features, there are thirteen numerical or numerically encoded features. We can run Principal Component Analysis procedure on this dataset and calculate which resulting features cover how much of the data set. After applying cumulative sum on results we receive plot which tells us, that first 6 components after PCA explains over 99% of variance in this data set.

As you can see, despite Bike Sharing Demand dataset is rather simple, it allows to do some data exploration and checking some of "common sense" guess about human behavior. Will this data work well with machine learning context? I don't know yet, but maybe I will have time to check it. If you would like to look at my step by step analysis you can check my github repository with EDA Jupyter notebook. I hope you will enjoy it!

Monday, April 14, 2014

Book review: Doing Data Science


Second book on my way to become "Data Scientist" is Doing Data Science: Straight Talk from the Frontline by Cathy O'Neil and Rachel Schutt. Here is short chapters description:
  • In first chapter authors try to deal with definition of "Data Science", "Big Data" and "Data Scientist" based on academia and industry experience and historical approach. They also polemics about data science as hype, extension of statistics of actual science.
  • Second chapter is dedicated to "big data" talk, and explanation why exploratory data analysis is important.
  • Chapter three introduces reader to three basics algorithms used in data science: linear regression, k nearest neighbors and k means. There are nice examples of actual use in GNU R.
  • In chapter four we will met naive Bayes in context of spam filters. We will also see example of building such filter with bash.
  • Chapter five is dedicated to logistic regression. Linear regression is explained on advertising company and example in R is provided. 
  • Chapter six - here comes data with time. Two examples are discussed - first is recommendation engine based on time data, and second is market stocks price analysis. 
  • Chapter seven covers feature selection problem. Influence of feature selection is discussed with connection with decision trees and random forests. Kaggle competitions idea is also discussed in this chapter. 
  • Chapter eight is dedicated to recommendation engines. It covers various methodologies for creating such engine and there is also simple example written in Python. 
  • Chapter nine is dedicated to visualization. I had very mixed feelings about this chapter. First part of it is dedicated to various visualization "installations" across different places. I could hardly find anything useful here. But on the other hand, second part shows how simple but clever visualizations could impact day to day search for fraud in credit card related business.
  • Chapter ten is dedicated to various problems and definitions among social networks.
  • Chapter eleven describes problems when it comes to determine causality. Cause and effect elements may be obvious in some situations, but in some they might be impossible to distinguish and measure.
  • Chapter twelve dedicated to data science in epidemiology describes fundamental problems with working with medical data. On the first approach good looking research could produce results whatever you like (positive or negative). Authors showed that there is surprisingly low interest in designing proper models to avoid such mistakes.
  • Chapter thirteen deals with data leakage. Author mentioned that, especially in data science competitions, often prepaired data has additional information that could product model that will work nicely with train set, but will work badly with real life data.
  • Chapter fourteen covers Hadoop and some elements of its ecosystem. Authors points to problem of Big Data: "Why do we need such solutions like Hadoop?" and "How to use them properly?".
  • In chapter fifteen students are summarizing they experience with learning data science during course which is combined in previous chapters.
  • Last chapter is dedicated to discussion about future of data science. Authors try to summarize their predictions about future and give some hints to aspiring data scientist how to look further.
In general, I have pretty mixed feelings about this book. I will start with negative aspects of this lecture. In my opinion it is pretty chaotic. Chapters are not connected. They are written based on materials form different persons representing different aspects of data science. It is more like bunch of different stories packed together than one big story. Also, some chapters have some mathematical formulas, which lead to nowhere. And often general discussion leads to place when reader anticipates great "finale" but it is not there.

On the other hand, if you treat this book like complementary material to "How to become data scientist" seminar on your university it could work pretty well. It covers many stories and problems which are usually visible only for persons doing actual business with data. And it gives good look and feel about doing data science. 

If you are expecting this book to be a handbook you will dislike it entirely. It is not even close to handbook. But if you like to read stories from actual data scientist working in business you will like this book.

Monday, February 24, 2014

Mining Titanic with Python at Kaggle

If you are looking for newbie competition on Kaggle, you should focus on Titanic: Machine Learning from Disaster. It is one among Getting Started competition category. It is indented to allow competitor getting familiar with submission system, basic data analyze, and tools like Excel, Python and R. Since I'm new in data science, I will try to show you my approach to this competition step by step, using Python with Pandas module.

First of all, some initial imports:
 import pandas as pd  
 import numpy as np  
We are importing pandas for general work and numpy for one function.

File which is interesting for us is called "train.csv". After quick examination we can use column called "PassengerId" as Pandas data frame index:
 trainData = pd.read_csv("train.csv")  
 trainData = trainData.set_index("PassengerId")  

Lets see some basic properties:
 trainData.describe()  
 numericalColumns = trainData.describe().columns  
 correlationMatrix = trainData[numericalColumns].corr()  
 correlationMatrixSurvived = correlationMatrix["Survived"]  
 trainData[numericalColumns].hist()  
Describe function should give summary (mean, standard deviation, minimum, maximum and quantiles) about numerical columns in our data frame. Those are: "Survived", "Pclass", "Age", "SibSp", "Parch" and "Fare". So lets take those columns and calculate correlation (Pearson method) with "Survived" column. We receive two results which are indicating two promising columns: "Pclass" (-0.338481) and "Fare" (0.257307). This is going well with intuition, because we expected that rich people will somehow organize them better in terms in survival. Maybe they are more ruthless? Lets see histograms for those numerical columns (last line in above code):
Lets examine other, non numerical data. We have "Name", "Sex", "Ticket", "Cabin" and "Embarked". "Name" column consist unique names of passengers. We will skip it, because we can't possibly correlate survival with this value. Of course, based on name we could determine if someone is V.I.P. of some kind, and then use it in model. But we don't have this data, so without additional research it is worthless. Then we have "Ticket" and "Cabin" values. Some of them are unique, some are not. Those values are also differently encoded. Knowing the layout of cabins on ship and encoding system we also could try to work with columns, but we also don't have such data. The same situation is with "Embarked". If persons are placed on Titanic with respect of place of embark, this could have major impact on survivability. People in front parts of ship had significantly less time to react when crash occurred. But we also don't know how embark could affect placement on ship. Maybe I will try to examine this column in later posts. Last non-numerical column is most interesting one. It is called "Sex" (my blog will probably be banned in UK for using this magical keyword). Sex column describes gender of persons on Titanic. Using simple value_counts function we can determine structure of this column (as well as others mentioned earlier):
 trainData["Sex"].value_counts()  
We receive "male 577" and "female 314". So that's was the problem. Male and female are strings, not numerical values. But we can easily change it:
 trainData["Sex"][trainData["Sex"] == "male"] = 0  
 trainData["Sex"][trainData["Sex"] == "female"] = 1  
 trainData["Sex"] = trainData["Sex"].astype(int)  

What can we expect from gender and chances to survive on Titanic? As we remember form movie "Titanic", persons which were organizing evacuation from sinking ship, said that women and children should go first to emergency escape boats. I'm not sure how accurate movie was, but they showed that in fact there was mostly women on emergency boats. And we have some hints about that when we calculate correlation again: 0.543351. This is nicer result than with "Pclass" and "Fare".

First idea which comes to my mind is to build classification tree with first split according to "Sex" parameter. Lets calculate overall survival ratio and then survival ratio for males and females:
 totalSurvivalRatio = trainData["Survived"].value_counts()[1] / float(trainData["Survived"].count())  
 totalDeathRatio = 1 - totalSurvivalRatio  
 maleSurvivalRatio = trainData[trainData["Sex"]==0]["Survived"].value_counts()[1] / float(trainData[trainData["Sex"]==0]["Survived"].count())  
 maleDeathRatio = 1 - maleSurvivalRatio  
 femaleSurvivalRatio = trainData[trainData["Sex"]==1]["Survived"].value_counts()[1] / float(trainData[trainData["Sex"]==1]["Survived"].count())  
 femaleDeathRatio = 1 - femaleSurvivalRatio  
So we have: totalSurvivalRatio = 0.38383838383838381, maleSurvivalRatio =
0.18890814558058924 and femaleSurvivalRatio = 0.7420382165605095. It looks like being male or female does somehow affect your chances on sinking ship. Of course, this make sense if there is time to evacuate and limited seats in boats. When everything would be more rapid and placed in different conditions favors may be reversed.

To be more formal, lest calculate Information Gain. In simple words, information gain is amount of entropy reduction after splitting group by parameter. To calculate it, we need to estimate total entropy in set, then entropy in subsets which emerge after division main set by "Sex" and subtract them with respect of probability of being in a group. Again, simple code in Python:
 males = len(trainData[trainData["Sex"]==0].index)  
 females = len(trainData[trainData["Sex"]==1].index)  
 persons = males + females  
 totalEntropy = - totalDeathRatio * np.log2(totalDeathRatio) - totalSurvivalRatio * np.log2(totalSurvivalRatio)  
 maleEntropy = - maleDeathRatio * np.log2(maleDeathRatio) - maleSurvivalRatio * np.log2(maleSurvivalRatio)  
 femaleEntropy = - femaleDeathRatio * np.log2(femaleDeathRatio) - femaleSurvivalRatio * np.log2(femaleSurvivalRatio)  
 informationGainSex = totalEntropy - ((float(males)/persons) * maleEntropy + (float(females)/persons)* femaleEntropy)  
So, total entropy (calculated for value "Survived") equals 0.96070790187564692. Since we expect it to be between 0 (no entropy, perfect pure set) and 1 (maximally impure set), it means more or less that there are similar numbers of persons who died, and survived. So plain guessing will not be much effective. After division for males and females, entropy looks different: maleEntropy = 0.69918178912084072 and femaleEntropy = 0.8236550739295192. And informationGainSex = 0.21766010666061419 which is pretty nice result. To be precise, we should calculate information gain for other columns, but I'm ignoring this step for this tutorial.

Since we calculated male and female survival ratio, we can use them to classification that female are surviving and males are dying. Such simple classification should give you a 0.76555 accuracy on Kaggle public leaderboard. I will try to achieve better result in next part of this tutorial.

Monday, February 17, 2014

Book review: Data Science for Business

Since I'm interested in data science but I'm newbie into this field of research, I decided to read some introductory book dedicated to this topic. Firs book which I read was Data Science for Business written by Foster Provost and Tom Fawcett. Below are my descriptions of each main chapter of this book:
  • First chapter of this book is dedicated to overall definition of data science, big data and similar.
  • Second chapter introduces "canonical data mining tasks".
  • Chapter 3 shows first steps with supervised segmentation and decision trees.
  • Next chapter adds linear regression, support vector machine and logistic regression. 
  • Chapter 5 - in my opinion most useful - defines overfitting. Authors shows examples how one can hit the overfitting problem, but also shows how to avoid it and deal with potential problems.
  • Chapter 6 introduces additional data science tools: similarity, neighbors and clustering methods.
  • Chapter 7 focuses on aspects strictly related to applying earlier mentioned tools to business - expected profit. This well written chapter shows that there is almost always second bottom, apart of pure data tools - business bottom.
  • Great data scientist, at some point has to show his results and hypothesis to stakeholders. He can use lots of complicated mathematical formulas, but also can use simple plots with additional information to nicely visualize his ideas. Chapter 8 describes some fundamental "curves" which are often used in data science.
  • In chapter 9, authors describe Bayes' rule and discuss its advantages and disadvantages.
  • Chapter 10 is dedicated to "text mining". Authors know, that they just scratch top layer of this issue. But on the other hand, reader can find here some basics ideas how to work with text and how to start researching different methods.
  • Final evaluation of example problem which was used through this book is done in chapter 11. 
  • Chapter 12 discuss other techniques with approaching analytical tasks: co-occurrence and associations (example usage: determining item which are bough together). Profiling, link prediction and data reduction is also discussed with nice example of Netflix Prize. Authors also clearly explain why ensemble of models could give better results in some cases.
  • In chapter 13, authors shows how to think about data science in business context, but also points how to work as data scientist in business environment.
  • Last chapter is dedicated to overall summary. Authors gives hints how we should ask data science related questions and how to think in general about data science.
Lecture of this book was very satisfying. I wasn't hit by enormous quantity of new definitions, equations and examples. For newbie in data science, reading this book chapter after chapter is like going step by step after your mentor. Using one main example during whole book was great idea. Reader can observe different techniques and problems related to them, applied on the same business situation. Also, business awareness is raised from chapter to chapter. I recommend this book either to data scientist wannabe or to "suit" who want to hire some geeks to examine business possibilities which gathered data.

Actually I can't say anything bad about this book. Of course, I would just love to see a complementary handbook with code in Python or R, but I guess that there are plenty of such books.

Sunday, February 9, 2014

Test your data science skills - Kaggle competitions

I recently started directing my interests towards Data Science. I started reading related books and doing related MOOCs. It is very interesting area of science, which has particularity good application in business. But dry learning from books and from MOOCs could be unrelated to real life problems and also could be boring at some point. For programming, solution is easy: coding competitions. Here is list of many pages dedicated more or less to coding competitions. But how about data science?

Luckily, there is one website which is hosting data science competitions: kaggle.com. There are several categories of contest organized on kaggle:
  • Featured: often complex problems but heavily sponsored in terms of prizes. Organized by big companies.
  • Masters: limited access competitions. You have to receive access to Master Tier of kaggle by achieving great results in previous competitions.
  • Recruiting: competitions dedicated for recruitment
  • Prospect: competitions without leader boards. Goal of them is usually to explore various data sets and discuss results among other kagglers.
  • Research: problems related to strict research areas.
  • Playground: interesting problems which are solved for fun.
  • Getting Started: tutorials
By the time I'm writing this post, there are 14 competitions (4 of them are long lasting tutorial competitions), with 334 K$ prize pool.

How are those competitions working? When you register for each competition, you receive access to train and test data and some additional information. Train data has known value that you have to predict. Test data is actual data on which you will test your model. After running model on test data, you submit your results to automated test web application and receive accuracy results. Those results will build public leader board. It is called public because only part of your results are taken into calculation of your score. Other part will be taken into account after closing the competition and based on it, final leader board will be constructed. By this split, it is harder to overfit model by examining results.

So how to start with data science competitions? I recommend starting with Titanic competition. This competition has nice tutorials in Excel, Python and R, and looks quite easy, at least at beginning. 

Anyway, I wish you GL & HF. See you on leader boards!

Friday, January 31, 2014

Specializations at Coursera - new quality among MOOCs?

I am big fan of MOOCs. Idea of preparing lectures, quizzes and assessments for people to be accessed through web browser, mostly for free is greatly generous. And its not just idea. There are foundations, universities and companies which are actually doing this. With bigger or smaller success but still. Currently my favorite site which is offering MOOCs is coursera.org.

How are MOOC working? Lets assume for example, that you want to learn Python. You maybe have some books, you maybe read some examples across web, and wrote some simple script. But you also might feel kind of lost. You don't know what parts are important, and maybe you don't have idea how to change newly acquired knowledge into practice. And here comes MOOCs. MOOC is Massive Online Open Courseware. It means that, if you find interesting MOOC, it will provide series of lectures (often video lectures), some simple quizzes after each part of material and bigger homework after larger segment. At the end, there usually is exam, maybe project or something similar, designed to test learned skills. So, back to Python. If you check site mentioned earlier, and answer the question "What would you like to learn about?" with "python" you will find three upcoming courses at the day this post was written. Those are: "An Introduction to Interactive Programming in Python", "High Performance Scientific Computing" and "Learn to Program: The Fundamentals", prepared by Rice University, University of Washington and University of Toronto. All those courses are free, and represents different approaches to topic with different difficulties and time frames.

But what to do if we want to learn something more deeper? We can search for complementary MOOCs over many different sites. There is an problem - often complementary course offers large part of it as introduction which might double the information that you already learned, thus lead to waste of time. As solution to this problem, Coursera prepared "Specializations". Specializations are series of courses which are covering bigger ideas. For me, their recent offer which is dedicated to Data Science is perfect. I'm interesting in data science since some time, but just recently I started to look for related MOOCs. And I found mentioned specialization. Don't worry about prices listed there. Those prices are for official printed and signed certificate of accomplishment issued by university and Coursera. If you don't need such certificate, you may enter those courses for free, and if you complete them, you will receive simple free PDF version. Not to mention knowledge ;)

I'm very excited by this idea, and I hope it will be widely adopted by other sites offering MOOCs. Damn, I actually can't imagine, how general learning will look like in twenty years form now :)