RFC: A separate, read-optimized projection of OpenMRS clinical data

For me, I would want this to remain a module for Platform 2 compatibility and consider moving to core in Platform 3, though I am ever mindful of this brilliant post. At the very least, I think we need more rapid development of the query store functionality than is likely tolerable in core.

This is definitely an interesting question and I’d really value the input of people not named Daniel Kayiwa or Ian Bacher on this one. CQRS is generally a pattern applied to systems built on event sourcing and while we’ve been playing around with improvements to events in OpenMRS, the core data model itself definitely is not. Eventual consistency has some real tradeoffs not only in processing overhead (as the write representation is translated the the query representation) and there is a (probably noticeable in real-world scenarios!) gap between a write and the query store being updated.

That said, I think there’s a very strong argument to be made that some kind of query store can enable a lot of functionality that the community has asked for. One way to think about this is that the query store effectively stores “flattened” records of OMRS transactional data and this module gives us the ability to “plug-in” different generated representations. This won’t solve everything we want flattened data for (i.e., it’s probably not the right representation for large-scale analytic queries), but it does seem a valuable representation to back things like patient flags, CDSS rules, calculated obs, etc.

I think Postgres only makes sense in this role if we also once again work on support Postgres as the transactional data store.

The other obvious storage to consider is Infinispan which is clusterable, supports vector stores, and is part of core since TRUNK-6302 / 2.8.0.

MariaDB also supports a vector storage (as of 11.8, which landed last year), but which I’m not sure we’ve tested with OMRS and isn’t as widely-used as pgvector. As near as I can tell, MySQL’s vector storage landed only in their paid products.

I favour purpose-driven, custom representations. Maybe it makes sense for the FHIR module to serialize FHIR representations (which could be one way to speed-up the FHIR2 module), but I don’t think FHIR is always the right representation for everything, especially AI RAG use-cases. Basically, if we want to provide FHIR serialization, I’d vote for this to live in the FHIR2 module.

Granular privilege management is something that we as a product have historically done quite poorly and I do think this is something we need to address to get to a MVP, but I assume for now that queries are largely patient scoped?