DEPTH-DVM Subprocessor Watch for Privacy Policies

A demonstrator that can identify and classify subprocessors and potential related problems in a privacy policy document: That’s what Coreon and Wult agreed on in course of their meeting for the DEPTH-DVM EU project.

Why are subprocessors important? A subprocessor is a third party that a service provider engages to process personal data on its client’s behalf, such as a cloud host or analytics vendor. Every subprocessor extends the chain of parties handling personal data. If just one of them stores data in a third country without adequate safeguards, the entire arrangement may no longer be GDPR-compliant. And since subprocessor lists change frequently, continuous monitoring is essential.

Multilingual Knowledge Graphs for Compliance

Wult is a Danish company offering a compliance platform designed to check policies in the context of digital advertising. About two years ago, they approached Coreon with the intention of integrating the Multilingual Knowledge Graph into their solutions. Since they receive policies in various languages, mainly European, they recognized the need to map different European privacy laws in a knowledge graph as a foundation for further tools.

That is exactly what we did over the past year: we built a Multilingual Knowledge Graph focused on the domain of privacy law, comprising about 700 concepts and 4,500 terms. This resource is valuable not only for translators and legal practitioners, but also for improving decision-making processes in privacy law.

Consider, for instance, testing an LLM’s ability to correctly classify a policy as containing—or missing—disqualifying, critical, or other relevant statements. When comparing pure LLM reasoning to LLM reasoning supported by the knowledge graph, the latter showed clear improvements in reasoning capability.

Knowledge Graphs support LLM Reasoning

One example illustrates this well: a policy states that personal data may be transferred outside the EU to Japan, the US, and China. The LLM-only approach merely notes that data transfers to China without supplementary measures are concerning, whereas the LLM-with-KG approach correctly identifies that this disqualifies the policy.

The reason is that the LLM was prompted to consult the Multilingual Knowledge Graph for additional information and found that China is denoted as a “High-Risk Third Country,” which the graph labels “Disqualifying” if no supplementary measures are described in the policy.

Infographic Subprocessor Watch

Demonstrator presented at Language Intelligence and teKom

These experiments form the basis of the demonstrator, which will be presented at the Language Intelligence conference in Vienna this November. Interested visitors can also stop by our booth at tekom the week before, also in November.

The goal is to move from purely static data to an interactive UI capable of simulating “trigger” events, such as the discovery of a new subprocessor. The system will use a rule engine to flag specific risks, such as non-EU data destinations, among others. The final output will be a one-page report in HTML or PDF format, using a traffic-light classification system to indicate compliance status.

Gefördert vom

Bundesministerium für Bildung und Forschung

Supported by the Eurostars programme, co-funded by the European Union
and national funding bodies.

Eurostars
Carina Obster
Carina Obster

Carina Obster studied Translation/Chinese Studies in Munich and Data Science in Vienna. As part of her role at Coreon GmbH, she has been working on Knowledge Graphs and their potential applications for over two years.