Check database quality before starting a Biology IA
Audit provenance, sampling effort and definitions so your database analysis supports a defensible biological question.

Inspect the source behind the spreadsheet
Start with the organization that produced the dataset, its collection purpose and its documentation. Record the version or access date, download location and any conditions of use. A large table is not automatically reliable biological evidence. You need to understand how observations became the rows you will analyse.
Choose a question the database can actually address. If you want to investigate species richness, check whether the records describe complete surveys or opportunistic sightings. Those collection methods can produce very different interpretations of an apparent absence.
Define a row and a variable
Write what one row represents: an individual, survey, location, population or time interval. Inspect the variable definitions, units and coding. A zero may mean an observed absence, an unrecorded value or a category code. Missing data should remain identifiable until you have a justified way to handle them.
In an illustrative bird-survey dataset, 'duration' could mean time spent observing at a site, while 'distance' describes a moving route. Clarifying those definitions changes which records can be compared and which factors might influence the response.
Audit coverage and sampling effort
Check where, when and how often data were collected. Repeated observations from a small number of accessible places may dominate a region-wide dataset. Experienced observers may identify more species than beginners, while longer surveys may detect more species regardless of habitat quality.
For a bird-richness question, compare effort variables before interpreting site differences. You might restrict records to comparable survey types or analyse effort explicitly. Explain the consequence of that decision: narrowing records can improve comparability while reducing the range to which your conclusion applies.
Create a transparent cleaning record
Preserve the downloaded data and work on a separate copy. Record each cleaning rule, its rationale and the number of observations affected. Check duplicate records, impossible values and inconsistent categories against the documentation. A surprising biological value should be investigated before being discarded.
For example, two rows with identical location and date may represent duplicate uploads or separate surveys. Treating them as duplicates without checking the survey identifier can remove real evidence. Make decisions using provenance and definitions rather than a preference for a smoother graph.
Test one small analysis
Inspect a manageable subset and calculate the response exactly as your question requires. Plot it against the explanatory variable and identify where comparability remains weak. Choose statistical methods with your teacher's guidance and explain their assumptions. Many rows do not automatically create many independent observations.
- Document inclusion and exclusion rules before the final analysis.
- Check whether records share sites, observers or collection conditions.
- Keep measurement units consistent across sources or years.
- Avoid joining datasets until location and time definitions match.
State what the dataset can establish
Use the conclusion to describe the relationship your design supports. An observational association may have several biological or sampling explanations. Identify the limitation most likely to change the interpretation and suggest a feasible improvement, such as a more consistent survey method or a restricted comparison. Bring the data audit and pilot plot to your teacher before investing in the full investigation.
Your next-step checklist
Checklist choices stay in this browser.
Official references
This is an original practical guide from IBvia. The examples are illustrative; your subject guide, assessment year and school instructions determine the requirements.

