Good data management is not a step you bolt on before analysis. It runs alongside the study from protocol to lock. Before the first participant is enrolled, we translate your protocol into a collection instrument and a set of validation rules. During conduct, we clean incoming records, raise and resolve queries, and code adverse events and medications. At the end, we reconcile, freeze, and hand off a documented, CDISC-aware dataset. Framing the engagement this way matters because most of the cost of a messy database is paid at the end, when errors surface during analysis and every fix has to be traced back through unversioned spreadsheets. Building the data management plan and the edit checks up front is what keeps that from happening.
Every engagement starts with a data management plan (DMP). This is the document that tells everyone, including a future auditor, exactly how data will be captured, cleaned, coded, stored, and locked. Ours specifies the study variables and their formats, the data dictionary, the coding conventions, the validation logic, the query workflow, and the roles responsible for each. A DMP is also increasingly a funding requirement: the NIH Data Management and Sharing policy and most institutional review boards now expect a written plan, and we prepare DMPs that satisfy those requirements as a standalone deliverable or as part of broader grant methodology support. A plan that is written once and followed consistently is the single strongest predictor of a clean database at lock.
We design the case report form (CRF), on paper or as an electronic data capture (eCRF) form, so that each field maps cleanly to a protocol variable and to your planned analysis. Where you already run a platform, we build inside it; where you do not, we work in open, well-supported systems such as REDCap and OpenClinica that are appropriate for academic and investigator-led research. Good form design prevents bad data at the source: sensible field types, drop-downs with controlled terminology instead of free text, skip logic, and range limits. On top of the form we specify edit checks, the automated validation rules that flag an out-of-range lab value, an impossible date sequence, or a missing required field the moment it is entered rather than months later.
Data cleaning, validation, and query management
Once data starts flowing, the work becomes data validation and query management. We run the edit checks, review the data for patterns the checks cannot catch, and raise discrepancy queries against records that look wrong. Every query is tracked from open to resolution in a documented log, so at any point you can see how many are outstanding and who owns them. Where the protocol calls for it, we support source data verification, checking entered values against source documents. This cleaning discipline is what separates a dataset that is merely complete from one that is correct, and it is the same rigor we bring to standalone data analysis service work and to survey and questionnaire datasets that need cleaning before they can be analysed.
Medical coding and reconciliation
Adverse events, medical history, and concomitant medications have to be coded to a standard dictionary so they can be grouped and analysed consistently. We code adverse events and medical history to MedDRA and medications to WHODrug, following your versioning and auto-coding conventions, and we reconcile coded terms against the verbatim source and, where relevant, against safety data. Consistent coding is what lets a reviewer count how many participants had a given class of event rather than sifting hundreds of near-duplicate free-text strings. For teams whose primary need is safety literature rather than trial data, this pairs naturally with our pharmacovigilance literature screening work.
CDISC-aware datasets and database lock
When collection is complete, we clean the final discrepancies, confirm coding and reconciliation are closed, and prepare for database lock, the point at which the dataset is frozen for analysis. We structure the delivered data to be CDISC-aware: organised along SDTM (Study Data Tabulation Model) principles with controlled terminology, so that if your study later needs formal SDTM and ADaM datasets for a regulatory submission, the groundwork is already in place. At lock you receive the analysis-ready dataset, the data dictionary, the full query and audit trail, and lock documentation. From there our biostatisticians can pick the analysis up directly, since statistical analysis and reporting is delivered by the same team, or you can hand the locked dataset to your own statistician.
Data management for grants, dissertations, and academic trials
Not every study is a registered clinical trial, and our clinical data management services are scoped for the reality of academic research. A PhD candidate cleaning a longitudinal dataset, a research group running a single-site registry, and an investigator-led trial preparing for a first regulatory conversation all need the same fundamentals: a written plan, a validated database, disciplined cleaning, and a defensible audit trail. What changes between them is the depth, not the standard. The same data dictionary discipline that stops a doctoral dataset filling with unusable free text is what keeps a registry analysable five years in.
Standards, quality, and audit readiness
Data handling for clinical research is judged against Good Clinical Practice and the ALCOA data-integrity principles: data should be attributable, legible, contemporaneous, original, and accurate. We work to those principles and design our documentation so an engagement is inspection-ready, with awareness of 21 CFR Part 11 expectations for electronic records where your platform and study require it. This is the same documentation-first discipline behind our regulatory literature reviews: the goal is that every value in the final dataset can be traced back to its source through a complete, dated trail. We are transparent about scope. We are a specialist academic and research data management team, not a full-service contract research organisation, and we will tell you plainly at the quote stage which parts of your study we are the right fit for.
Most datasets that reach a statistician in poor shape share the same handful of failures, and every one of them is cheaper to prevent than to repair. Free-text fields that should have been drop-downs produce dozens of spellings of the same value. Dates stored as text break every calculation that depends on them. Missing data goes unrecorded, so nobody can tell a true zero from an unanswered question. Silent duplication of participants inflates the sample. Undocumented mid-study changes to how a variable was collected quietly bias the results. Our up-front data dictionary, typed fields, controlled terminology, and edit checks close each of these off at entry, and our query management log catches the rest before database lock rather than during analysis, when a single fix can mean re-running the entire statistical plan.
Timelines are set by your study, not by a queue. The data management plan and the database build are discrete pieces of work with their own dates, and you get both in the written quote before you commit to anything. Cleaning and query management run alongside enrolment for as long as the study is collecting data, so that phase lasts as long as your study does. Database lock follows the last data point, once the final queries are closed and coding is reconciled. If you are joining us mid-study, the first date we commit to is the written assessment of your existing database, delivered before any cleaning begins, so you know what you are buying before the larger engagement starts.
What we need from you, and how we scope
To quote accurately we need three things: your protocol or study plan, the platform you use or intend to use, and your timeline. From those we scope the engagement to what your study actually requires. A doctoral project may need only a data management plan, a REDCap build, and a cleaned dataset; a registry may add ongoing cleaning and periodic exports; a small trial may run the full path through medical coding, reconciliation, and a formal database lock. Pricing is fixed per scope and held constant regardless of your location, so the number you approve is the number you pay. Where your need is really statistical rather than data-handling, we will point you to the right service instead of overselling this one, whether that is our biostatistics team or funder-ready methodology writing.
Your engagement is led by a named methodologist, and the delivered dataset is checked by our Director of Biostatistics before handoff, so there is real, credentialed accountability on every project rather than an anonymous queue. Any materials and data you provide remain your exclusive property at all times, we execute a mutual non-disclosure agreement within 24 hours of request, and we accept purchase orders for institutional work. Every engagement runs against a defined written scope agreed before work begins. If an auditor or reviewer later questions the data handling, cleaning, or documentation, we revise it at no charge. Tell us your study design, your platform, and your timeline, and a named methodologist will reply with a scoped plan and a fixed price.