Forecasting Infectious Disease Spread with Web Data

using Wikipedia to forecast infectious disease outbreaks — Incorporating real-time, anonymized data from Wikipedia and other novel sources of information is aiding efforts to forecast and respond to emerging outbreak.

Just as you might turn to Twitter or Facebook for a pulse on what's happening around you, researchers involved in an infectious disease computational modeling project are turning to anonymized social media and other publicly available Web data to improve their ability to forecast emerging outbreaks and develop tools that can help health officials as they respond.

Mining Wikipedia Data

"When it comes to infectious disease forecasting, getting ahead of the curve is problematic because data from official public health sources is retrospective," says Irene Eckstrand of the National Institutes of Health, which funds the project, called Models of Infectious Disease Agent Study (MIDAS). "Incorporating real-time, anonymized data from social media and other Web sources into disease modeling tools may be helpful, but it also presents challenges."

Latest Videos From

Watch full video here:

To help evaluate the Web's potential for improving infectious disease forecasting efforts, MIDAS researcher Sara Del Valle of Los Alamos National Laboratory conducted proof-of-concept experiments involving data that Wikipedia releases hourly to any interested party. Del Valle's research group built models based on the page view histories of disease-related Wikipedia pages in seven languages. The scientists tested the new models against their other models, which rely on official health data reported from countries using those languages. By comparing the outcomes of the different modeling approaches, the Los Alamos team concluded that the Wikipedia-based modeling results for flu and dengue fever performed better than those for other diseases.

TOPICS