Overview
Impact Data Intelligence is a framework for transforming unstructured information about natural-hazard events into structured, reusable datasets, developed for the OCEANIDS project (EU). It collects information from sources such as news websites, Wikipedia, and scientific publications, and extracts both quantitative and qualitative impacts using LLMs.
Challenge
Impact information is scattered across sources and usually appears as narrative text. The system needs to handle differences in terminology, dates, locations, languages, and levels of detail while identifying duplicate reports of the same event. It also needs to remain independent of any single LLM provider.
Methodology
The pipeline handles HTML and PDF ingestion, text preprocessing, location extraction, translation, summarization, and structured extraction through LLM function calling. Extracted events are validated, standardized, deduplicated, aggregated, and georeferenced using Nominatim. PostgreSQL, SQLAlchemy, and PostGIS provide storage, while Redis and Celery support asynchronous processing. A manually labelled dataset provides a basis for evaluating and comparing LLM performance..