This topic will serve as the central place for planning and tracking the Metadata Extraction Tool project throughout my OpenMRS Fellowship.
The goal of this project is to develop a Java-based utility that reads metadata from a live OpenMRS instance and exports it into Initializer compatible CSV files, making it easier to replicate existing OpenMRS configurations across different deployments.
Objectives
Define the metadata domains to be supported.
Design a modular and extensible extraction architecture.
Implement metadata exporters domain by domain.
Generate Initializer-compatible CSV output.
Ensure correctness through unit and integration testing.
Document the architecture and usage for future contributors.
Development Plan
The project will be completed incrementally:
Finalize the extraction scope with mentors.
Create Jira epics and implementation tasks.
Implement each metadata domain individually.
Add tests for each exporter.
Improve documentation and gather community feedback.
Project Tracking
I’ll be sharing the Jira board, implementation tasks, and design updates here as they become available, so the community can easily follow progress and provide feedback.
Repository:
I welcome suggestions, questions, and feedback throughout the project. Thanks in advance for your support!
@bawanthathilan - great initiative. We’ve started a few things like this in the past at PIH but never really taken them through to completion for various reasons.
Here is something I threw together somewhat recently as the start of an effort to move one of our implementations to Initializer, where my main initial goal was to just get their concepts and concept-related dependencies into code, but where I was planning for this to be something we could extend to other domains as necessary:
For Concepts specifically, one might question why I bothered with the Concepts domain when OCL is generally the preferred route, and the reason was because the starting dictionary I was working with had evolved since the earliest days of OpenMRS, and many of the concepts as-is would fail validation if re-saved or imported into a new system. Exporting Concepts as-is into the concepts domain allowed us to quickly iterate on identifying issues and cleaning them up prior to the more expensive and permanent process of putting them into OCL. This is also why I added some other tools into this repo to help along the way (eg. identify if concepts were used by analyzing foreign key references and other usages, etc)
Keep in mind when building a tool like this that many of the domains are quite customizable. For example, with Locations - some might choose to manage the tags associated with each location as additional columns in the location domain, others might prefer to use the locationtagmaps domain. Concepts is like this especially, and another complication is that many dictionaries may have inherited additional concept names or mappings that they don’t need or want to preserve. So it is not always a one-size fits all situation.
Thanks Mike , really appreciate you sharing this and the context behind it the ConceptExporter approach.
That point about domain customization makes a lot of sense too, especially with Locations and Concepts having multiple valid ways to model the same thing. @wikumc and I will keep that in mind so the tool doesn’t assume a one size fits all approach, maybe by allowing some configurability in how each domain gets exported.
Hi everyone! I’m creating a draft wiki page for the Metadata Extraction Tool. It’s still a work in progress, but I’d appreciate any feedback or suggestions. @wikumc@jayasanka@ibacher@dkayiwa
The core domains are complete. and we’re currently working on supporting the remaining domains. However, we still don’t have a UI that community members can use to test the tool
How about we prioritize the frontend integration and start a discussion on a frontend mockup. That would allow other community members to test the module and provide feedback. Before that I think we should merge Wikum’s PR first
Please let me know your thoughts, suggestions, or anything you think could be improved. Your feedback will help us refine the UI before we move forward with the implementation.
This is a really thoughtful design. Fantastic work @bawanthathilan!
Few things I’d change:
The UI currently shows endpoint info, like which endpoint got triggered and where stuff was fetched from. As a user I don’t really care about the internals, so I’d drop that.
Can you help me understand the usecase for downloading old builds? These files are meant to be committed to a content package, so they end up in version history anyway. Don’t really see the point of keeping them version controlled here too. That said, if we really want the last build and the backend supports it, we can show it. I’d just keep the latest build prominently and collapse the history. That’s what 90% of users want.
Probably a good idea to add a select all option when picking domains.
Something to keep in mind: there’s no retention policy for the zip files, so those could get deleted manually from the directory. In that case the download throws a 410 or something. So either we show a clear error on download, or somehow communicate to the user that these aren’t guaranteed to be available. Not a big thing tho.
Stuff like export logs and connect to a server, since we don’t offer them yet I’d just drop them. Not sure about domain settings. I couldn’t find an endpoint supporting domain settings.
Also, retire is an internal technical term for soft delete, users don’t need to know that. Just call it Delete here and don’t list deleted packages at all.
These files aren’t “meant to be committed to a content package”, these files are or potentially are literally the content package or at least they should be treated that way. (That is the goal!)
We do not enforce that users are storing these files stored in version control. This tool could very well be the version control, so the storage isn’t necessarily redundant.
My main piece of feedback here is that things are too broad. It is very unlikely that I want to export “Concepts”. I want to export “All Concepts” or “These 10 concepts” or “The concepts associated with this form”. I think the later works if I select a form, but the UI hides that and doesn’t give me an obvious way of checking what is or isn’t in any particular package.
FWIW “As a user, I want to be able to visualize the contents of a content package before I download it” is a clear user requirement we should support.
I’d drop the “build” language, as I don’t think that will be understandable to users. Maybe we call “builds” “snapshots” or “versions” or something? (OCL uses “versions”, so that’s at least something in the broader ecosystem using that terminology).
Also, just a comment on the API, but something we’re calling a “build” should probably not be run on a GET request. GET requests are intended to leave server state unchanged and a “build” that creates a new file does not do that. Data retrieval in response to other HTTP verbs is a standard part of HTTP operations.
Thanks, @ibacher and @jayasanka for pointing these out. Yes, we should include a select all option for concepts. And yes, there may be some keywords that users are not familiar with, so we should consider making those clearer.
@ibacher I agree that a build must not run on a GET, and that’s how the API is laid out. The only way a build is ever created is POST /packages/{uuid}/build