Scientific, research, and laboratory software

Biotechnology Research Workflow Platform Requirements

A requirements guide for biotechnology teams connecting protocols, samples, instruments, analysis, decisions, and governed research-data reuse.

Published by · Fact-checked by OpenAI Codex research review · Published · 1240 words

Define the scientific decision before choosing a platform

Biotechnology research software should preserve how evidence was produced, interpreted, challenged, and reused. Begin with the research questions and operating boundary rather than a feature list for an electronic notebook, laboratory system, data lake, or project tracker. Identify the programs, assay types, organisms or materials, instruments, collaborators, computational work, quality expectations, and decisions the first release must support.

Observe a representative experiment from hypothesis and approved protocol through material receipt, preparation, execution, raw output, processing, quality review, interpretation, and downstream decision. Include a failed run, protocol deviation, contaminated or insufficient sample, repeated analysis, corrected metadata, unavailable instrument, and departing collaborator. The laboratory system requirements checklist covers laboratory operations broadly; this guide focuses on exploratory research provenance and collaboration.

Model protocols, materials, runs, and results separately

Use durable identities for study or program, hypothesis, protocol and version, experiment, run, batch, material, sample, container, aliquot, reagent lot, instrument, method, raw file, transformation, observation, derived result, quality decision, interpretation, and publication output. Not every team needs every entity, but a single flexible “experiment” document rarely preserves lineage when methods and materials are reused across programs.

Separate planned work from execution. A protocol describes intended steps and parameters; an execution records what occurred, including deviations and environmental context. Preserve parent-child material relationships, quantities and units, custody, location, freeze-thaw or passage history where meaningful, consumption, and disposition. Record which protocol version, instrument configuration, software version, reference data, and parameter set produced each output so later analysis does not rely on filenames and memory.

Preserve provenance through computational analysis

Treat raw data as immutable evidence and transformations as versioned, reproducible work. Record inputs, checksums, code or workflow version, environment or container, parameters, reference datasets, operator or service identity, timestamps, logs, outputs, and failure state. Derived data may be regenerated only if the required inputs and execution context remain available. A chart copied into a presentation is not an adequate research record.

The NIST Research Data Framework organizes research-data concerns across envisioning, planning, generation or acquisition, processing or analysis, sharing or reuse, and preservation or disposal. Use that lifecycle to assign ownership and acceptance evidence. It does not prescribe one architecture, and the platform should accommodate approved discipline-specific repositories and computational environments instead of forcing every dataset into one database.

Design collaboration around rights and scientific context

Define organization, program, project, dataset, and item-level access with explicit roles and purpose. Collaborators may need to contribute samples, run a method, review selected outputs, or receive a released dataset without seeing unrelated programs, identities, inventions, or contractual information. Model embargo, publication, licensing, material-transfer, consent, export, and sponsor constraints as governed decisions rather than free-text warnings that users can bypass.

Comments and tasks should reference stable records and versions. Preserve who requested a change, what evidence was reviewed, the resulting decision, and whether later edits invalidate it. Support export in usable, documented formats with identifiers and provenance. Avoid building collaboration around downloadable spreadsheets that detach measurements from units, methods, corrections, and access conditions, then return as an untraceable source of truth.

Plan data sharing, retention, and privacy early

The NIH Data Management and Sharing Policy overview expects covered investigators to plan and budget for management and sharing and to follow an approved plan. Applicability depends on funding and research context. Capture the approved sharing plan, repository, timing, metadata, restrictions, responsible role, and evidence so obligations can be operated rather than rediscovered near publication.

Classify human, genomic, proprietary, export-controlled, confidential, and safety-relevant data with qualified institutional input. Minimize collection, separate direct identifiers, enforce approved use, log consequential access, and establish retention and disposition. If the system supports regulated clinical investigations, the FDA electronic-systems guidance provides relevant expectations, but exploratory software should not claim blanket compliance merely because it has timestamps and signatures.

Validate scientific integrity and operational recovery

Test identity, permissions, protocol versioning, unit conversion, precision, timezones, instrument import, duplicate handling, partial files, reruns, lineage traversal, corrections, export, archival retrieval, and restoration. Use representative large and unusual datasets. Reconcile sample and file counts, hashes, relationships, and key scientific values during migration. A successful row count cannot prove that a transformed dataset retains meaning.

Define backup and recovery objectives by research consequence, then restore in a clean environment and trace a result back to raw evidence. Monitor unavailable instruments, stalled imports, orphaned files, failed analyses, permission changes, unusually broad exports, storage growth, and overdue reviews without putting sensitive science into telemetry. Preserve a documented exit path so the institution can continue research if a vendor, collaborator, or custom developer is no longer available.

Migrate evidence without flattening its scientific meaning

Inventory notebooks, shared drives, instrument workstations, databases, analysis environments, sample systems, repositories, and personal working folders. Select representative old, large, sensitive, incomplete, and unusual records. Map identities, units, controlled terms, protocol versions, material lineage, files, checksums, transformations, comments, and access conditions. A migration that moves final PDFs while abandoning raw files, context, and correction history may create a cleaner interface and weaker science.

Define what will be migrated, linked in place, archived read-only, or deliberately retired. Rehearse extraction and load with the production pipeline, preserve source identifiers and transformation logic, and reconcile counts plus meaningful scientific relationships. Require researchers and data stewards to trace selected conclusions before and after migration. Quarantine uncertain matches rather than guessing, and keep a documented route back to the original source until acceptance and retention decisions are complete.

Operate vocabularies, integrations, and change over time

Assign ownership for controlled terms, units, reference datasets, protocol templates, instrument connectors, computational environments, access groups, retention rules, and repository mappings. Vocabulary changes need identifiers, definitions, effective dates, mappings, and review; replacing a label globally can corrupt historical interpretation. Treat connector and analysis updates as scientific changes when they can alter parsing, precision, normalization, classification, or derived results.

Release changes through representative validation datasets with expected outcomes, lineage checks, permission tests, performance measures, and rollback criteria. Monitor unknown terms, failed imports, changed file shapes, orphaned relationships, stale reference data, unavailable computational dependencies, and growing manual correction queues. Fund ongoing stewardship and researcher support. A platform cannot preserve reproducibility if no one owns the definitions, integrations, and environments on which reproducibility depends.

Provide search and discovery without detaching results from scientific context. Index approved identifiers, descriptions, protocol and material relationships, provenance, and access conditions while respecting embargoes and project boundaries. A search result should explain why a record matches, which version is shown, and whether related evidence is incomplete. Derived recommendations or similarity features must remain advisory, disclose their data boundary, and never relabel a scientific conclusion or expose another program's confidential work through snippets, counts, or suggestions.

Establish a retirement process for instruments, methods, repositories, and computational dependencies. Before removal, identify affected active work, preserved evidence, export formats, replacement mappings, validation needs, and long-term retrieval responsibilities. Keep sufficient documentation and executable context to interpret important historical results. A license ending or workstation failing should not make an organization's earlier research impossible to inspect, reproduce where required, or transfer to an approved successor environment.

Pilot one reproducible research path

Select a bounded assay or workflow with representative materials, instrument output, analysis, quality review, collaboration, and sharing. Require the team to repeat the analysis from preserved inputs, explain every transformation, identify the responsible protocol version, restrict a collaborator correctly, correct metadata without erasing history, export an approved dataset, and restore the record after a simulated failure. Compare time, error, completeness, and reproducibility against the current process.

Before requesting an estimate, prepare sample workflows, protocols, material and data types, instruments, computational tools, repository obligations, collaboration agreements, sensitivity classifications, volume projections, difficult exceptions, migration sources, and acceptance evidence. Submit them through the project brief for a phased solution review, or use quick contact to decide whether integration, configuration, or focused custom research software best fits the boundary.

Authoritative references

Related software planning guides

Explore custom software development