I was asked a question the other day: “How many biodiversity data points do you think there are in Australia?” There’s nothing like an interesting question to get your mind working.
The biggest collection of biodiversity data points in Australia is in the Atlas of Living Australia (ALA). As I was writing this, there were over 184 million records in the ALA. So let’s start there…

The Atlas contains data from a wide range of data providers, as you can see on the ALA’s dashboard (https://dashboard.ala.org.au/). When you look through the top ten of these data providers, you’ll see some interesting - and very large - groups of records that are worth thinking through:

- Citizen science charges into the lead with birders around Australia contributing records through the Birdata platform and the NSW bird atlas,
- State Government Departments - with NSW winning the race through their Bionet platform but chased closely by South Australia, the Northern Territory,
- Our collections community - through Australasian Virtual Herbarium (AVH) and Online Zoological Collections of Australian Museums (OZCAM) they provide the combined data from the collections in our Museums and Herbaria around Australia,
- Our global infrastructure - including the Australian Antarctic Data Centre (AADC*), GBIF and OBIS - who have collated data from all over the world (and a subset of which goes into the ALA).
* Note: Maybe the AADC belongs in the government departments section…
So that adds up to 80.54 million records - less than half of the records in the Atlas - so there’s then a long tail of other providers, with none being more than 600,000 records. There’s a lot of different other providers in there but as a whole you can probably say most of the records are from similar types to the top 10.
When you look at these records though - you probably have some biodiversity data that’s not being delivered to the ALA from these organisations; data on species that are under taxonomic review, or embargoed because they are Restricted Access Species Data (RASD, see the framework for more information on those reasons). So maybe there’s another 10% of records being withheld from the Atlas - there’s another, say 18 million records still locked away by the existing providers.
And then there are projects that aren’t in the ALA - yet, at least. These include programs like the fisheries bycatch programs around Australia - a lot of that doesn’t end up in this ecosystem of systems (with an exception of NSW - where there’s around 2 million records coming into the ALA at the moment through that top listing for the ALA datasets). Let’s assume that there’s similar fisheries bycatch data for each jurisdiction in Australia (sorry ACT, not including you here) and that’s another 14 million records. Let’s assume there are a few other datasets out there in those providers that aren’t in either, so being conservative, let’s just round that to 15 million.
Let’s also look at who’s not showing up in the ALA listings? Who’s missing from this list, and how big of an unknown set of data is out there? The ones that stick out to me straight away are industry data, research data and local government data.
Industry data
Industry data is something that’s really interesting to look at, and it’s something that probably can’t be separated cleanly from the consultants that do the work collecting that information. Across Australia, there are hundreds of environmental consulting companies that collect biodiversity data and provide it to their clients, such as mining companies seeking approval to mine areas, offshore renewable energy companies seeking to establish infrastructure and a wide range of companies representing any industry that has an impact on the environment. A lot of this data ends up in reports that are used in the environmental impact assessment process, often called the “grey literature”. Way back in 2005, when Gaia Resources was still operating from my house, I wrote a paper about this grey literature in Western Australia, and in our experience, this still is an issue across the country.

The reports I managed to source for that 2005 paper were really the tip of the iceberg about what was out there…
In WA, there have been some valiant efforts to make this data available, from the organisations themselves (such as work we did with the Great Victoria Desert Biodiversity Trust recently). Government is doing a lot, too: The Department of Water and Environmental Regulation has done a lot of work making reports and data packages available through their projects like the Index of Biodiversity Surveys for Assessment and the Index of Marine Surveys for Assessments, as well as their successor system, Environment Online. Some of this data also finds its way into our state aggregator, the Department of Biodiversity, Conservation and Attraction’s Dandjoo system, where you can explore all the datasets through their datasets dashboard. Our work with the Western Australian Biodiversity Science Institute on the Shared Environmental Analytics Framework is also unlocking a range of data to be used, at least collaboratively between industry participants. It’s been really rewarding for Gaia Resources to be involved in so many of these projects, and through them there’s more biodiversity occurrence data becoming available - and there are similar initiatives occurring around the country.
So industry data - while there is a general perception that industry is sitting on a mountain of data like a dragon guarding their gold, this isn’t really the case. There are some really amazing detailed, local datasets around mining infrastructure in areas, but you might only be talking about 20-30 million records that aren’t already in the data ecosystem somewhere.
Research data
Researchers are often the holders of an amazing detailed dataset for their area of interest. They do have tremendous value in particularly narrow scopes - doing detailed research on a single population of an animal is going to be super useful for things like population viability and other ecological questions, but it’s probably only going to generate a small number of biodiversity occurrence records. Let’s add a few million records to the pot for researchers.
Local government
Our third tier of government in Australia is local government; this is where we have an interesting potential for biodiversity data. According to the Australian Local Government Association (ALGA), there are over 500 councils in Australia, which starts to add up when you think about it. If you were to assume that each of these held a hundred thousand biodiversity records, that’s 50 million records. But is that likely?
According to ALGA, some of the things that local governments do includes animal control, planning and development approval which all might be ways in which biodiversity occurrence data is captured. In our previous work with local governments around mosquito control (via the Atlas of Environmental Health), they were trapping hundreds of mosquito sites and capturing in some cases thousands of individuals - there is a huge potential for records just in the environmental health area (although these are almost certainly RASD!).
So maybe 50 million isn’t too much of an estimate from local governments after all.
What does all that add up to?

One of the things to consider is the way that ecological surveying practices are changing. As techniques like environmental DNA sampling, automated camera traps and machine learning for identification improve, there could be a significant uptick in the volume of records that comes in through platforms like Wildobs, which already holds over 18 million images across over 600 species (at the time of writing). This will mean more and more data coming into the various biodiversity platforms… so I fully expect this to be a rather conservative estimate (a bit like my old paper that definitely underestimated the amount of grey literature out there).
Of course, not all of this data is captured with the same processes, rigour and expertise; and that’s where the biggest challenge arises: what is the quality of this data? This is something that we’ll have to explore in another blog later, but let me just say this as a teaser: quality is in the eye of the beholder. More on that to come…
In the meantime, another 106 million records are out there somewhere. There’s a number of great programs like the ALA’s Data Mobilisation Program that is aiming to help get them moving. If you have data that you are sitting on and want to discuss how to get it moving, or you want to know how to access so much of this value that is out there, then give us a call on (08) 92277309, drop us a line via [email protected], or start a conversation with us on our social media channels.
Piers
Cover image credit Tanja Tepavac on Unsplash