Welcome to Data Science Hong Kong

Data science is starting to become embedded in Hong Kong. We are a community for all data scientists in Hong Kong — from beginners to multi-decade practitioners of BI, artificial intelligence and data warehouse design, and from students to professors — and their fellow travellers, from business, government and academia. We want to create an environment where data scientists can learn from each other and share their stories, and a community that non-data scientists can turn to when they want to understand more about what data science can do for them. We will organise monthly events such as unhackathons and lectures as well as hosting social media platforms.

Join us on our other social platforms to keep up-to-date with our activities and to become part of our community:


slack_icon

meetup_icon

FBicon

linked-in_icon

Un-hackathon #10

Our 10th Hackathon for Data Science: a full day of fun and working together on YOUR data science projects!

At this event attendees will have the chance to pitch their projects, or join other people’s. And in the beginning of the day we will host some fantastic industry specialists to share their experiences operating in the data science field.

Signup at: Eventbrite, Meetup, Facebook

The event will be held at the South China Morning Post offices at Times Square

 

Schedule of events:

9.30am – Arrive, registration
10am – Welcome
10.15am – Talks begin
11.30am – Pitch session, recruitment
12pm – Work on projects
5.30 pm – Present results of work session

Location:

SCMP: 20/f, Tower 1, Times Square, 1 Matheson St, Causeway Bay

Requirements:

Laptop / charger for those joining the coding
Prepared data, and projects pitches for the ones submitting projects
If presenting, send us your presentation slides ahead of time so we can prepare them.
50HKD in cash for the space rental

Recommendations for project submissions:

Send us your presentation slides! Drop a link to one of the organisers on Slack or another way. We want to minimise time spent switching laptops so we will run your slides from our pc.
Prepare data in advance as much as you can; spending the day cleaning or retrieving data won’t gather crowds of DS! Contact organisers if you need a data repository to share data with all your team members.
If the project is already underway, prepare an introduction to it so that people can join. If you’re presenting slides, send them to us before you arrive, make sure the task you propose is feasible during the time of the event, and describe the skills you expect your team to have: R or Python? AWS, Spark? etc.

For final presentations:

Start writing the final presentation right from the start and add elements little-by-little all day long. Articulate the reason you want to do the project, and the solution. Make it understandable to everyone.
If you wish, your work will be published on this website with your bio, name, etc.

Other details:

50 participants max
Food/drink: Only water, coffee and tea are provided. Attendees can order their own food to the venue, take a break to find a restaurant nearby or bring their own lunch.
Price: 50 HKD. We charge a fee to cover venue and food costs. We are a not-for-profit organisation and will aim to keep the costs of our events as low as possible to make it accessible to all.

Unhackathon #9 roundup

With the World Cup in Russia wrapping up on the same evening as our ninth un-hackathon, football was on our minds, and, with tongue in cheek, Data Science Hong Kong co-organiser Xavier put the question to our cohort of predictive modellers to find the winner, hours before the result was known.

IMG_20180715_165227
We were asked to build a predictive model for the world cup result, but our vote worked well enough. 

Getting down to more serious stuff, Houston Ho presented his company’s work on using machine learning to predict whether an employee is due to leave their position. His tool aims to give human resources teams a score for each employee based on a variety of characteristics. His model is achieving 80% accuracy and he says he can achieve more.

See his presentation below.

DSHK co-organiser Guy Freeman also presented about his new development offering a central repository for scraped data, using an open source philosophy. He showed the system’s potential by using a dataset of property transactions in Hong Kong spanning 20 years.

See his presentation below.

Group projects

Michael attracted the most interest among the group in his restaurant prediction model project. Using data from two restaurant booking systems, he aimed to predict how busy a restaurant would get using a machine learning model.

See the results from the group’s work in the slides below.

Forecasting Visitors of Restaurants

Visiting from Tokyo, Suzana Ilic brought exciting skills to the unhackathon, and decided to set her sights on unpicking hype in the crypto space. She was a bit camera shy so no video but you can see her slides below.

Quantifying Hype

Image from iOS

Morris Wong aimed to build an auto-tagging system for publishing, along the lines of taggernews. This tech could have wide-scale application when it’s up and running. See his presentation below.

Pocket exploration

That’s it for this month’s event. We will be announcing our next event shortly.

Unhackathon #8

Our 8th Hackathon for Data Science: a full day of fun and working together on YOUR data science projects!

At this event attendees will have the chance to pitch their projects, or join other people’s. And in the beginning of the day we will host some fantastic industry specialists to share their experiences operating in the data science field.

Signup at: Eventbrite, Meetup, Facebook

naked-hub

Schedule of events:

9.30am – Arrive, registration
10am – Welcome
10.15am – Talks begin
11.30am – Pitch session, recruitment
12pm – Work on projects
5.30 pm – Present results of work session

Location:
16F, 40-44 Bonham Strand, Sheung Wan, Hong Kong

Requirements:

Laptop / charger for those joining the coding
Prepared data, and projects pitches for the ones submitting projects
If presenting, send us your presentation slides ahead of time so we can prepare them
50HKD in cash for the space rental

Recommendations for project submissions:

Send us your presentation slides! Drop a link to one of the organisers on Slack or another way. We want to minimise time spent switching laptops so we will run your slides from our pc.
Prepare data in advance as much as you can; spending the day cleaning or retrieving data won’t gather crowds of DS! Contact organisers if you need a data repository to share data with all your team members.
If the project is already underway, prepare an introduction to it so that people can join. If you’re presenting slides, send them to us before you arrive, make sure the task you propose is feasible during the time of the event, and describe the skills you expect your team to have: R or Python? AWS, Spark? etc.

For final presentations:

Start writing the final presentation right from the start and add elements little-by-little all day long. Articulate the reason you want to do the project, and the solution. Make it understandable to everyone.
If you wish, your work will be published on this website with your bio, name, etc.

Other details:

50 participants max
Food/drink: Only water, coffee and tea are provided. Attendees can order their own food to the venue, take a break to find a restaurant nearby or bring their own lunch.
Price: 50 HKD. We charge a fee to cover venue and food costs. We are a not-for-profit organisation and will aim to keep the costs of our events as low as possible to make it accessible to all.

Unhackathon #7 round-up: making sense through data

What time do people rent share bikes in San Jose? Houston and a group of data scientists has looked at bike share data in California and made some curious obvservations at our April unhackathon.

We also heard from Nick Lam-wai who is building a database on Hong Kong’s budget, the blueprint of government spending and priorities. And Chris Choy, who was working with Nick also discovered how to take historical PDFs of the budget and read the tables into Nick’s database. Expect big things from this group.

Our second meet up at Accellerate in Sheung Wan started with a discussion of the  Catboost library by Daniil Chepenko, who explains its benefits over other methods such as random forest.

Catboost is a gradient boosting library for work on decision trees, developed by the Russian search engine Yandex, building on many years of development in this field.

See his presentation video below, and follow the slides here.

Projects

Willis sought to find out what makes a Kickstarter project work. He came to the hackathon with data from 2009-2017, and a trained model with 60% accuracy, up from 30% at the beginning of his work. Knowing whether a Kickstarter will succeed is a huge investment advantage, so watch the short videos to see how well he went.

Pitch:

Conclusion:

Elizabeth Briel and Ben Davis have been seeking new ways to tell the story of global warming’s effects on arctic sea ice, and came to the hackathon with data they wanted to turn into a song. See the results below.

Slides are here.

Pitch:

Conclusion:

Nick Lam-wai created a thorough database of the Hong Kong budget, turning it from a human readable collection of documents back into one ready for machine analysis.

Slides here.

Pitch:

Conclusion:

Overwatch strategies revealed with data science

Ram de Guzman presented this analysis of Overwatch team strategies using scraped data from Winston’s Lab (which gathers it directly from game videos). His insight revealed how the best teams in South Korea arranged their teams and fought.

In the video he describes the process of gathering his data, then shows in impressive visualisations how that data relates to actual game strategy.

Watch his talk at our 6th unhackathon in March here:

 

And you can follow his project here.

Data science news round up

Our tight-knit community of data scientist have shared a wealth of news and inspiring projects from around the web over the past couple of months. Here is a brief round up of the more interesting articles, and remember, you can join in on our slack group.

2-l-304106-unsplash

Millions of Chinese farmers reap benefits of huge crop experiment

An article that demonstrates the world changing potential of evidence based approaches to the world’s problems. For me, it’s also a reminder that it’s often not the latest buzzword or most glamourous topics that have the most impact.

Winning with Data Science

Next is an article examining the business and organisational side of data science. This is a topic that probably doesn’t get enough attention compared to the latest and coolest algorithm. It’s important for data scientists to take an interest in how organisations should adapt, if they don’t it will probably be decided by someone not qualified to make the decision!

nasa-43569-unsplash

What Comes After Deep Learning?

This article examines whether deep learning is actually a blind alley and considers what new approaches might be next for data science. Also a brief examination of the question of US vs China in the AI “arms race”.

‘Who’s Leading AI’ Isn’t the Intelligent Question

Our final article explores the much talked about question of whether the US or China is winning and why it’s not the right question to ask.

If you found any of these articles interesting then do come and join the discussion on our Slack group, where you will also find details of meetups. https://datasciencehk.slack.com/

April 15 Unhackathon #7

poster_7

We are organizing another Un-Hackathon on April 15th! You can sign up here! We have organized a number of talks and a day of collaborative, hands-on problem solving.

Details:

  • 9.30am – Arrive, registration
  • 10am – Welcome
  • 10.15am – Talks begin
  • 11.30am – Pitch session, recruitment
  • 12pm – Work on projects
  • 5.30 pm – Present results of work session

Location:

11F, 40-44 Bonham Strand, Sheung Wan, Hong Kong

Requirements:

Laptop and charger for those joining the coding.
Prepared data and project pitches for those submitting projects.
If presenting, send us your presentation slides ahead of time so we can prepare them.
50HKD in cash for admin and organisation.
Recommendations for project submissions:
Prepare data in advance as much as you can; spending the day cleaning or retrieving data won’t gather crowds of DS! Contact organisers if you need a data repository to share data with all your team members.
If the project is already underway, prepare an introduction to it so that people can join (if you’re presenting slides, send them to us before you arrive), and make sure the task you propose is feasible during the time of the event, and describe the skills you expect your team to have: R or Python? AWS, Spark? etc.

For final presentations:

Start writing the final presentation right from the start and add elements little by little all day long. Recall the context of the project and articulate the presentations to make it understandable by the non-initiated public around you.
If you wish your work will be published on the website datasciencehongkong.com with your bio, name, etc.

Other details:

50 participants max
Food / drink: Only water, coffee and snacks are provided. Attendees can order their own food to the venue, take a break to find a restaurant in Kennedy Town or bring their own lunch.
Price: 50 HKD. We charge a fee to cover costs. We are not a for-profit organisation and will aim to keep the costs of our events as low as possible to make it accessible to all.

March 18 Unhackathon #6

nh-sw-hk-1Data scientists: we are organising another Un-hackathon in our monthly series. There will be talks and a day of collaborative, hands-on problem solving.

Sign up on our Eventbrite page and stay up to date with upcoming events on our meetup page.

 

Details:

  • 9.30am – Arrive, registration
  • 10am – Welcome
  • 10.15am – Talks begin
  • 11.30am – Pitch session, recruitment
  • 12pm – Work on projects
  • 5.30 pm – Present results of work session

Location:

16F, 40-44 Bonham Strand, Sheung Wan, Hong Kong

Requirements:

Laptop and charger for those joining the coding.
Prepared data and project pitches for those submitting projects.
If presenting, send us your presentation slides ahead of time so we can prepare them.
50HKD in cash for admin and organisation.
Recommendations for project submissions:
Prepare data in advance as much as you can; spending the day cleaning or retrieving data won’t gather crowds of DS! Contact organisers if you need a data repository to share data with all your team members.
If the project is already underway, prepare an introduction to it so that people can join (if you’re presenting slides, send them to us before you arrive), and make sure the task you propose is feasible during the time of the event, and describe the skills you expect your team to have: R or Python? AWS, Spark? etc.

For final presentations:

Start writing the final presentation right from the start and add elements little by little all day long. Recall the context of the project and articulate the presentations to make it understandable by the non-initiated public around you.
If you wish your work will be published on the website datasciencehongkong.com with your bio, name, etc.

Other details:

50 participants max
Food / drink: Only water, coffee and snacks are provided. Attendees can order their own food to the venue, take a break to find a restaurant in Kennedy Town or bring their own lunch.
Price: 50 HKD. We charge a fee to cover costs. We are not a for-profit organisation and will aim to keep the costs of our events as low as possible to make it accessible to all.

Women in data science – WiDS 2018

The Stanford Women in Data Science conference 2018  is starting on March 6th at 1am Hong-Kong time

Live Broadcast

We encourage everyone to follow the broadcast here 

You can tweet using the hashtag #WiDS2018Q

Program

The program can be found here, we reproduce it here for convenience in HK time zone

1:00-1:10am: Opening Remarks: Margot Gerritsen, Senior Associate Dean and Director of ICME, Stanford University
1:10-1:30am: Welcome Address: Maria Klawe, President, Harvey Mudd College
1:30-2:05am: Keynote Address: Leda Braga, CEO, Systematica Investments
2:05-2:10am Regional Event Check-in
2:10-2:50am: Technical Vision Talks:
     2:10-2:30am Mala Anand, EVP, President, SAP Leonardo Data Analytics
     2:30-2:50am Lada Adamic, Research Scientist Manager, Facebook
2:50-3:10am: Morning break
3:10-3:15am: WiDS Datathon Winners Announced
3:15-3:55am: Technical Vision Talks:
     3:15-3:35am: Nathalie Henry Riche, Researcher, Microsoft Research
     3:35-3:55am: Daniela Witten, Associate Professor of Statistics and Biostatistics, University of Washington
3:55am-4:30am: Keynote Address: Latanya Sweeney, Professor of Government and Technology in Residence, Harvard University
4:30-6:00am:  Lunch and Breakouts (NO LIVESTREAM)
6:00-6:35am: Keynote Address: Jia Li, Head of Cloud R&D, Cloud AI, Google
6:35-7:15am Technical Vision Talks:
     6:35-6:55am: Bhavani Thuraisingham,
Professor of Computer Science and Executive
Director of Cyber Research and Education Institute, University of Texas at Dallas
     6:55-7:15am: Elena Grewal, Head of Data Science, Airbnb
7:15-7:30am  Afternoon break 

7:30-7:35am Regional event check-in
7:35-8:15am Career Panel moderated by Margot Gerritsen
Bhavani Thuraisingham 
 Professor of Computer Science and Executive
Director of Cyber Research and Education Institute, University of Texas at Dallas
     Ziya Ma,  Vice President of Software and Services Group and Director of Big Data Technologies, Intel Corporation
     Elena Grewal Head of Data Science, Airbnb
     Jennifer Prendki, Head of Data Science, Atlassian
8:15-8:55am: Technical Vision Talks
     8:15-8:35am: Risa Wechsler, Associate Professor of Physics, Stanford University
     8:35-8:55am: Dawn Woodard, Senior Data Science Manager of Maps, Uber
8:55-9:00am: Closing Remarks