ESG and AI Report: Designing a Reliable LLM Architecture

01.10.2026

Every year, dozens of Italian companies produce a sustainability report. It is a job that always follows the same structure – interviews, data collection, mapping in line with ESG standards – but the content varies completely from client to client. In short, it is a fairly repetitive process, so much so that it can be partially automated.

This was the background when one of our clients, an ESG consultancy firm, contacted us. For us, it represented something new: not an AI integration into a product, as we had done previously with the PrivateAI Suite, but the first time we had placed an LLM at the heart of a system’s architecture.

What you are about to read is an account of how we found our way around an unfamiliar area, the choices we made and, above all, why we made them.

We started with the assessment (thankfully)

The first step was to work with the client to define the main objectives and key points of the project. From there, we drew up an initial draft of the project, focusing on requirements, needs and possible directions for development.

We then held discussions with BI-REX, a Competence Centre specialising in innovation and applied research, with whom we entered into a partnership not too long ago. During these sessions, we validated the initial project idea and worked together to determine how to turn it into a concrete AI-based project.

The aim was to work out what to build and, above all, how to go about it – that is, which technologies to use, how to integrate AI into the project, and how to structure a solution that would be genuinely applicable to the client’s context.

The entire project development process was not straightforward. Each meeting had a specific issue to resolve: how to manage input data, how to design the conversational flow, how to coordinate the agents, and how to measure the quality of the output. However, the greatest contribution was not technical, but methodological: having a framework for thinking through the problem before rushing to solve it.

It was at this stage that we realised the value of the PoC (Proof of Concept), from which the project itself took shape: starting from that, rather than aiming straight for the finished product, allowed us to build things properly and to understand what wasn’t working.

How to become an ESG consultant

To gain a better understanding of the context in which the programme is used, the client shared with us material from sessions that had already taken place, including recordings, transcripts and notes. We analysed these to understand what was happening in those texts, even before considering how to replicate it.

And what we learnt could not have been gleaned from any written brief. An ESG consultant never turns up for an interview unprepared. They have a detailed, well-established framework in mind, and use it to guide the conversation.

The questions may seem open-ended, but they are anything but. Underneath lies a structure that is far more rigid than one might think: a hierarchy of typical responses, expected pathways, and indicators of anomalies that require further investigation. The consultant follows a precise sequence, decides when to delve deeper, and already knows what to expect – and what not to expect. Getting an LLM to do all this might seem straightforward. It isn’t.

The point, therefore, is not simply to obtain answers because, in a field as specialised as this, an unguided AI produces answers that are broad but of little practical use. To obtain relevant answers, we needed a guided process and specialised agents, each focusing on a specific ESG area. And it was precisely at this stage that we realised what would make the difference.

From unstructured data to structured data

In practical terms, that difference has taken the form of an architecture. After analysing the data and the client’s approach, together with the BI-REX consultants, we identified the real objective: to achieve the best possible output. From there, an architecture emerged that is designed to manage each step in a controlled manner, orchestrated by a specific agent.

And so we’ve arrived at a guided flow. The AI agent dynamically generates response options based on the context of the interview; the user selects the one that best matches their situation and can respond freely if none of them apply.

This might seem like a limitation compared to an AI that is free to converse. In fact, the opposite is true. A constrained model remains consistent because its focus is narrow; a guided user gets less tired and makes fewer mistakes. And the data you collect – the real output, given that it ends up in a report with legal standing – is of a much higher quality.

In applied AI, freedom is not a value in itself. Often, the best result comes from the narrowest scope. We arrived at this through trial and error: mistakes, adjustments and iterations, until we found the right approach.

Three observations we will take forward into our future projects

  1. Expectations of an LLM are almost always more optimistic than reality. Many people take it for granted that AI will know on its own when to be precise and when to cut corners. In fact, in complex contexts, a model left to its own devices produces plausible but largely unhelpful responses. And in a regulated field such as this, ‘largely unhelpful’ means unusable, because that data ends up in a document with legal standing. Aligning expectations early on with what the technology actually does is just as much a part of the job as writing code.
  2. A well-structured process from the outset improves the end result. We spent hours talking to the BI-REX consultants, and every hour spent on assessment saved us ten hours of technical iteration. The difference wasn’t simply adding AI to a process, but understanding how an ESG consultant thinks and translating that logic into architecture. Without that understanding from the outset, we would have built a tool that was quick but flawed.
  3. In regulated sectors, it is essential to design with change in mind. Standards and regulations are constantly evolving, with frequent updates. In this context, an AI system cannot be rigid. That is why we have structured our work so that every regulatory update is a configuration task, rather than a complete rewrite. This is a strategic architectural choice, as it allows us to do the work once and never have to do it again.

Where we are now

The collaboration with BI-REX concludes with a PoC, followed by an MVP that includes an interface hosting the model, which users can interact with.

We then carried out tests involving real interviews, in collaboration with the client. Every piece of feedback served as a starting point for iterating and refining our work. From there, the project took shape: a recognisable brand identity and a platform that is already available on the market. You can find all the details in the full case study.

Throughout the project, we certainly made some mistakes – especially at the start, when we were navigating uncharted territory. This did not discourage us; on the contrary, it showed us just how much the approach matters as much as the technology, perhaps even more so. This is exactly what we found when working with BI-REX: a method capable of guiding us through the most complex stages.

But there is one thing we take with us more than anything else: namely, that when we tackle the next project of this scale, we will always start by asking the right questions.