WorkServicesBlogTalksAboutBook a call
Life Sciences

Opening Africa's Medical Data Presents Both Pros and Cons

Originally published in The Yuan on June 30, 2022, as part of its open medical data series. Republished here by the author. Read the archived original.

The future of healthcare will be data-driven and built on open data. This will be for several reasons, including accessibility, creating opportunities for independent researchers, universities, and companies, improved transparency, and collaboration that enables contributors to take data and make actionable insights out of them.

This article will discuss the benefits of open data and shed some light on possible drawbacks. As with every innovation or technological advancement, one must pay close attention to all side effects and unintended consequences, especially since many of these drawbacks are not all that unique to healthcare.

With regards to open healthcare data, Africa remains underrepresented, and the best remedy for this is to further promote open science and open data. For modern healthcare innovations such as precision medicine, drug discovery, and others to become fully established within the continent, the right foundation for them must be put in place by making the necessary data available.

There are a few open science projects that are ongoing, including popular ones such as the H3ABionet project and Data First. Thanks to open-source communities and initiatives, the African continent is gradually being absorbed by the radical open-source movement in technology, which should have a positive effect on open science.

That said, here are some advantages to open data.

Pros

Innovations in research and technology: Data being contributed by and made accessible by many people will boost the speed of innovation. This offers easy access to healthcare data to a wide population of healthcare professionals, researchers, universities, and hospitals, which can enhance various areas of research such as machine learning (ML) for healthcare, drug discovery, precision medicine, and genetic studies.

Open data and open science also encourage collaborations between industries and researchers or universities, building otherwise unlikely connections and diverse research teams.

Open data reduce research costs: While making data readily accessible to a vast population, open data also reduce the research costs required for data gathering, processing and annotations, which benefits small research organizations that would otherwise be unable to afford to pay for the data they need to get started.

Reduced costs will also be a major plus for African researchers who are mostly underfunded and lack access to the data they need because they are hidden, unavailable, or too expensive to access. Open healthcare data can serve as a remedy to all these problems.

Better informed healthcare policies: Open data open up conversations and meaningful discourse around digital health policies. As more work is done within the healthcare space, more questions will be raised, and the need to lay down laws will arise. There are currently 42 African countries with known eHealth strategies. As more people gain access to health data, researchers and institutions will raise concerns, and advise and influence governments on how to make policies to benefit everyone. With better healthcare policies in place, healthcare providers, governments, and researchers will all be better prepared and informed of how to prevent or predict catastrophic events like pandemics.

Representation of African Data: Africa has the most genetically diverse population of any continent and it is where most human evolution occurred, so it only makes sense for it to be included in healthcare research. Even so, very few Africans have been involved in these studies to date. With open data that Africans are actively contributing to, it will become much easier to study this diversity properly. Also, open data will give Africans an opportunity to have more of a say in how their data are used in public databases.

Encourages interoperability in healthcare: Open data sharing stirs conversations about healthcare data interoperability. There must be some defined structure or standard for data to be easily merged or migrated across different platforms and organizations. This is because these data standards cannot be set in isolation. They require a collaborative effort, which open data and science communities can provide.

With these standardizations, the data also become more impactful since they will be easier to analyze and interpret. This in turn improves the quality of available data, removing inconsistencies and ensuring that researchers have access to good quality data for their research and ML datasets.

Other benefits of open data to consider include contributions to achieving United Nations' sustainable development goals, inclusivity, fairness in AI, and bridging data gaps.

Cons

Like every other technological innovation, open data do have some downsides, and a few of them will be discussed here.

Data privacy issues and the mosaic effect: For open data to be safe, there must be discussions and actions about the best way to de-identify and anonymize the datasets beyond simply removing a person's name and location. The mosaic effect talks about how open-source datasets, when combined, can be traced back to individual people. One dataset in isolation will not cause any danger, but if it is combined with other ones, it could pose an increased risk, especially when made accessible to public users who may have malicious intentions. This combined information could reveal sensitive information, not just about individual persons but also vulnerable groups.

Available data could be biased: Bias occurs when a distribution of a particular dataset population only represents a few people and does not truly reflect an entire population. For example, with open datasets, the data made available could be provided by people who have access to certain facilities such as the internet, education, or healthcare facilities, which could leave out data about people who lack access to these things. Another thing to put into perspective is how in these datasets, certain countries such as Nigeria or South Africa could have very good representation considering their large populations with many interested contributors. However, this will not be inclusive of countries with few to no contributors.

Also, algorithmic bias is not restricted simply to race. There are also gender inequalities that exist. Prediction models for cardiovascular diseases that have been trained on predominantly male datasets, e.g., cannot accurately predict this disease in women, since there are different patterns of expression in men as compared to women. An algorithm trained on a predominantly male dataset would not be able to accurately diagnose heart attacks in women.

Security: Discussing data privacy and the mosaic effect also raises the need to make public health datasets more secure. This then raises the need for data security guidelines to protect vulnerable populations whose data will be made easily accessible to the public. These security measures must be quick and able to readily adapt to the constantly changing data landscape.

Other cons that should be considered include sustainability costs, governance, missing information, and incorrect use of data.

Conclusion

Though there are significant downsides, these are still outweighed by the benefits of open public health data. With increased participation in open science and open-source-related activities, the acceptance of open data represents a step in the right direction if the African continent is to seek to provide better solutions to health-related problems, take giant leaps in health research, and better explore and enjoy the benefits of modern medicine.

Contributions and collaborations in open data will stir conversations and encourage partnerships between research institutions, healthcare companies, universities, and hospitals, and educate the public on the benefits of sharing their data. African data have thus far been underrepresented, difficult to access or completely unavailable. Embracing open data would represent a significant step toward solving these problems.

On this page