In this article
Imagine two files describing the same synthetic examination. One reports a thickness as 0.52; the other reports 520. Both parse successfully. Both contain ordinary floating-point numbers. If one uses millimetres and the other micrometres, they agree. If the pipeline treats them as interchangeable numbers, the model sees a thousandfold difference.
This is the kind of problem that makes research engineering extend far beyond training code. Before a model can learn anything useful, the system has to preserve what its inputs mean. Afterwards, we need to know which inputs and transformations produced a result.
I work on these boundaries in CorneaForge, an on-premise platform for ophthalmic data. The examples here are synthetic. The architecture reflects code and design documents reviewed in September 2026, with the relevant components rechecked in October. The model activation workflow described below is a design contract with local components; it should not be read as a claim that the complete workflow is deployed or that a model has demonstrated clinical benefit.
Keep the original before interpreting it
Our fictional examination, EXAM-DEMO-001, arrives with a measurement file and a map. The first useful action is to retain the received objects and their identities. That gives later computations something stable to refer to.
This raw layer is often called Bronze. The name describes its purpose, not its quality. A corrupted file can be faithfully preserved. An authentic file can still contain an invalid measurement. Storage answers “what did we receive?”; validation answers different questions.
Suppose a device retransmits the same object after a network interruption. Receiving it twice should not silently create two examinations. The ingestion code needs an identity rule that recognizes the repeated input, then handles it consistently. This property is called idempotence: repeating an operation has the same intended effect as performing it once.
That is narrower than promising that every event will be processed exactly once under every failure. Identity can be incomplete, an export can change, and two files can belong to one acquisition. Those cases need explicit rules. A filename alone is rarely a sufficient explanation of why two records are the same observation.
Turn values into measurements
The next layer, Silver, gives the retained inputs an explicit structure. A value belongs to a particular acquisition, device, eye, time, quantity, and unit. Where available, it also carries a quality state and a reference to its source.
Our example might become:
acquisition: EXAM-DEMO-001
quantity: thickness
value: 0.52
unit: mm
origin: native device export
parser_version: demo-v1
The second file can map to the same canonical unit by dividing 520 µm by 1000. The original representation remains available. We have made a transformation, so that transformation has to be identifiable too.
Missing values deserve the same care. If the source has no thickness measurement, replacing it with zero creates a measured zero that never existed. It also changes what a downstream model can learn. Absence, rejection by a quality rule, and a legitimate numerical zero are different states.
Some quantities require more than unit conversion. Corneal data can involve polar coordinates, surface interpolation, fitted reference surfaces, or Zernike coefficients. I distinguish a native device value from a quantity recomputed by our code and from an image rendered for display. Similar names do not make these objects equivalent. A display map may incorporate interpolation and a colour scale; neither is an additional measurement.
Freeze the question you are studying
A structured database still is not a research dataset. A study needs a population, an observation unit, selected variables, labels, inclusion rules, and a construction date. Gold is the name used here for data assembled for research.
The observation unit matters immediately. Does one row represent an acquisition, an eye, or a person? Joining several devices without answering that question can multiply rows and quietly alter the weight of individual observations. The builder has to make its row definition visible before anyone calculates a metric.
CorneaForge includes a builder that writes research tables as Parquet and makes them available through DuckDB. Parquet supplies a column-oriented storage format; it does not decide whether a cohort is scientifically appropriate. That responsibility remains in the selection and transformation code. Apache Parquet’s documentation describes the storage format and its intended use.
Freezing a dataset gives “run the same experiment again” a concrete meaning. A path named latest cannot supply that meaning if new examinations or corrected values appear beneath it. A manifest can record a dataset version, a build identity, source references, and the transformations used to construct it.
A manifest need not copy every image. It can point to retained objects, provided those references really identify stable contents and remain available. Recording a digest helps detect a change; it does not restore a deleted file.
There is also a database-level problem during construction. A long extraction must define which committed state it reads. PostgreSQL’s default Read Committed isolation can expose different committed data to successive queries within one transaction, whereas Repeatable Read uses a stable transaction snapshot. Choosing and verifying a suitable extraction strategy is part of the builder’s job. A database snapshot during extraction and a preserved research dataset afterwards solve related but distinct problems. PostgreSQL documents these guarantees.
Research and inference take different paths
Once the data have an explicit meaning, the architecture branches.
The research path selects a cohort, separates development from evaluation, trains candidates, and records the evidence used to compare them. In ophthalmology, the two eyes and repeated visits of one person can be related. Randomly splitting individual rows may place closely related observations on both sides of an evaluation. The split policy has to match the intended question; patient grouping is one important part of that design.
The output of training is a candidate. Its existence does not establish that it is suitable for a particular use.
An inference path, by contrast, takes a new observation through the transformations required by an already selected model. It does not retrain a model for every examination, and it does not necessarily rebuild the research tables. What it needs is a compatible input contract and an explicitly chosen model package.
This distinction avoids a common source of accidental coupling: a research script writes a new weights file, and an application happens to read that location. A successful training run should not decide a deployment merely through its output path.
- BronzePreserve the source and its identity.
- SilverDecode, validate, normalize units.
Frozen dataset → patient-separated evaluation → candidate model.
Compatible new input → approved model version → attributable output.
This is an architectural explanation. A model candidate still needs validation and approval before activation. A research snapshot is not rebuilt for every inference request.
A model package includes the meaning of its inputs
Consider two pipelines that both emit a column called surface_rms. One computes a root-mean-square residual in micrometres. The other divides that residual by the fitted area and emits micrometres per square millimetre.
The column count is unchanged. The name is unchanged. The number has changed meaning.
This is why a model package needs more than weights. It needs the ordered feature schema, units, preprocessing version, and relevant handling of missing values. If a calibrator or decision policy is used, those belong to the compatible package as well. A raw score, a calibrated probability, and a thresholded decision are separate outputs with separate assumptions.
The lifecycle I am developing separates candidate creation, evaluation, explicit activation, and rollback. Activation should expose a complete compatible bundle to readers. A rollback must restore that bundle, including the input contract; replacing only the weights can recreate the same incompatibility in reverse.
These are engineering requirements whose enforcement has to be checked in code and runtime. A diagram with an approval box is not an approval mechanism. This broader dependence of ML behavior on surrounding components is also the subject of Sculley and colleagues’ Hidden Technical Debt in Machine Learning Systems.
Failure states carry information
Now interrupt the read of our synthetic examination. The original object is stored, but the worker cannot retrieve it.
An empty feature row would conceal what happened. The useful result is a visible, retryable read failure. When access returns, a worker can retry the transformation and record the outcome.
That case differs from a successful lookup proving that a required object is absent. Both differ from reading bytes successfully and discovering that they cannot be parsed. Retrying, requesting a corrected export, and declaring an input unavailable are different responses. Valid earlier results should not be erased merely because a later read failed.
In the concrete stack, MinIO retains objects, PostgreSQL holds structured data and state, FastAPI exposes request interfaces, and specialized workers perform background tasks. Systemd supervises services. Each component earns its place by carrying a responsibility; the stack itself is not the research contribution.
I want to be able to follow an output backwards: from its transformation version to a structured measurement, and from that measurement to the received object. That same trail helps answer a more practical question when something breaks: which step should we repeat, and which evidence should we preserve?
Try the boundary yourself
An extractor fixes a unit conversion without changing a column name. Can the existing model keep using its output?
Not automatically. Check the contract used during training, identify which data and models used the previous conversion, version the correction, and evaluate the compatible chain. Fixing the arithmetic and authorizing a model for use are separate decisions.
The database has gained new rows. Can yesterday’s study still be reproduced?
Only if its population, source contents, transformations, and evaluation choices remain identifiable and recoverable. Re-running the same query against today’s database is a new extraction, even when the query text is identical.