Skip to content
Pre-publication draft. This Trust Center is prepared for peer review before public launch.
Business Continuity and Disaster Recovery

Business Continuity and Disaster Recovery

Policy · v2026.06 · Owner: Security Officer / Engineering Lead · Effective: 2026-06-30 · Reviewed: 2026-06-30 · Next review: 2027-06-30

Bioscope Foundry maintains a documented contingency plan to restore the MSO operating platform, and the practices that depend on it, after a disruption. PHI resides exclusively in Foundry’s FHIR service (R4); recovery is therefore built on AWS regional and multi-region capabilities, Foundry’s FHIR service’s export and import procedures, and infrastructure-as-code rebuild automation. The plan is tested at least annually.

1. Purpose & scope

This Contingency Plan establishes procedures to recover Bioscope Foundry’s operations following a disruption resulting from a natural disaster, political disturbance, man-made disaster, external human threat, internal malicious activity, infrastructure-provider outage, or any other event that interrupts normal operations. It applies to Foundry’s operating platform (conductor on AWS), the agent control plane (foundry-master-agent and the Forge Agent fleet), Foundry-owned tooling, identity systems (the identity provider), and the work-site arrangements supporting Foundry’s workforce.

HIPAA reference: this Contingency Plan satisfies the HIPAA Security Rule’s contingency-plan standard at 45 CFR § 164.308(a)(7), including its required implementation specifications: data backup plan, disaster recovery plan, emergency mode operation plan, testing and revision procedures, and applications-and-data criticality analysis.

NIST reference: the plan is informed by NIST SP 800-34 (“Contingency Planning Guide for Federal Information Systems”) and is consistent with the Foundry security program’s broader alignment with the NIST cybersecurity guidance set.

2. Policy statements

Bioscope Foundry policy requires that:

(a) A plan and process for business continuity and disaster recovery (BCDR), including the backup and recovery of systems and data, is defined and documented.

(b) BCDR is simulated and tested at least once a year. Metrics are measured and identified recovery improvements are filed and tracked.

(c) Security controls and HIPAA requirements are maintained during all BCDR activities.

(d) Recovery time and recovery point objectives are defined for the operating platform and the PHI store, and are reviewed annually.

(e) When an engagement with a practice ends, Foundry executes the 90-day clean offboarding commitment, including return or destruction of PHI in accordance with the BAA.

3. BCDR objectives

The following objectives are established for this plan:

  1. Maximize the effectiveness of contingency operations through an established plan that consists of:
    • Notification / Activation phase: to detect and assess damage and to activate the plan;
    • Recovery phase: to restore temporary operations and recover damage done to the original systems;
    • Reconstitution phase: to restore full processing capabilities to normal operations.
  2. Identify the activities, resources, and procedures needed to carry out Foundry’s processing requirements during prolonged interruptions to normal operations.
  3. Identify and define the impact of interruptions to Foundry systems and to the practices Foundry supports.
  4. Assign responsibilities to designated personnel and provide guidance for recovering Foundry during prolonged periods of interruption.
  5. Ensure coordination with other Foundry staff who will participate in the contingency planning strategies.
  6. Ensure coordination with external points of contact and vendors (AWS, the identity provider, and any subprocessors handling PHI under BAA) who will participate in the contingency planning strategies.

4. Recovery time and recovery point objectives

Foundry defines distinct RTO and RPO targets for each tier of system. Targets are reviewed annually and after each tabletop or technical test.

System tierExample componentsRTORPORecovery posture
PHI storeFoundry’s FHIR service (R4) datastoreBeing establishedBeing establishedFHIR-service export to encrypted managed object storage; a continuous export schedule and cross-region replication are being established as part of the resilience buildout. Customer-managed encryption keys.
Operating platformconductor (services + SSR web)Being establishedBeing establishedContainer images and infrastructure-as-code in source control; production environment is reproducible from code; configuration is in secrets-managed stores; rebuild from scratch into an alternate region.
Agent control planefoundry-master-agent (Atlas, orchestrator, dashboard, policy)Being establishedBeing establishedRust workspace pinned by Nix flake; reproducible bootstrap. The control plane can be brought back without re-derived PHI access; Forge Agents resume only after Atlas is healthy.
Identitythe identity provider (workforce + providers)Being established≤ 1 hourManaged by the identity provider; recovery follows the provider’s continuity SLAs. Foundry maintains a documented break-glass account access path.
Non-critical systemsInternal dashboards, analytics, support tooling≤ 1 business day≤ 24 hoursRestored after the critical tier is online.

The targets above describe Foundry’s planning posture for the operating platform and are not contractual SLAs unless reflected in a Master Services Agreement or BAA. RTO is the time between the start of an outage and the resumption of service; RPO is the maximum tolerable data loss measured in time.

5. Critical and non-critical systems

Foundry classifies systems for DR purposes into:

  1. Critical systems: host production application services, the FHIR PHI store, the agent control plane, or are required for the functioning of those systems. If unavailable, they affect the integrity of data or the ability of practices to operate, and must be restored, or have a restoration begun, immediately upon becoming unavailable.
  2. Non-critical systems: all systems not classified as critical. They may affect performance or convenience but do not prevent critical systems from functioning or being accessed appropriately. Restored at a lower priority than critical systems.

The criticality assignment is documented in Foundry’s asset inventory and reviewed annually.

6. Line of succession

The following order of succession ensures decision-making authority for the Contingency Plan is uninterrupted:

  1. Security Officer: authority for the entire Contingency Plan; responsible for executing its procedures and ensuring the safety of personnel.
  2. Engineering Lead: responsible for the recovery of Foundry technical environments.
  3. Chief Executive Officer: assumes authority for the entire Contingency Plan if neither of the above can function in that role, or chooses to delegate to an alternative.

Succession contacts (recorded in the current org chart):

7. Response teams & responsibilities

The following teams have been developed and trained to respond to a contingency event affecting Foundry’s infrastructure and systems:

  1. IT: responsible for recovery of Foundry’s corporate environment (the identity provider, MDM-managed workstations, internal SaaS). Includes personnel responsible for daily IT operations and maintenance. Leader: Security Officer, reporting to the CEO.
  2. People & Facilities: responsible for ensuring physical safety of all Foundry personnel and environmental safety at each Foundry work site. Includes site leads where applicable. Leader: the Operations Lead, reporting to the CEO.
  3. Platform Engineering: responsible for restoring the operating platform (conductor), the agent control plane (foundry-master-agent), and their supporting cloud infrastructure. Also responsible for testing redeployments and assessing damage to the environment. Leader: Engineering Lead.
  4. Security: responsible for assessing and responding to all cybersecurity-related incidents according to the Incident Response policy and procedures. The Security team also assists the above teams in recovery as needed in non-cybersecurity events. Leader: Security Officer.

Each team lead maintains a local copy of this policy and the relevant team contact information (including a printed copy) so that the plan can be executed even if Internet access is unavailable. All executive leadership is informed of any contingency event.

8. General disaster recovery procedures

8.1 Notification and activation phase

This phase addresses the initial actions taken to detect and assess damage inflicted by a disruption to Foundry. Based on the assessment, the Contingency Plan may be activated by the Security Officer or the Engineering Lead. The Security Officer may also activate the plan in the event of a cyber disaster.

Notification sequence:

  • The first responder notifies the Operations Lead. All known information is relayed.
  • The Security Officer contacts the Response Teams and informs them of the event. The The Operations Lead or delegate begins assessment procedures.
  • The Security Officer notifies team members and directs them to complete the assessment procedures to determine the extent of damage and estimated recovery time. If damage assessment cannot be performed locally because of unsafe conditions, the Security Officer follows the alternate-assessment path below.
    • Damage assessment procedures. The Security Officer assesses damage logically, gains insight into whether the infrastructure is salvageable, and begins to formulate a recovery plan.
    • Alternate assessment procedures. Upon notification, the Security Officer follows the procedures for damage assessment with the Response Teams from a safe location.
  • The Contingency Plan is activated if one or more of the following criteria are met:
    • Foundry’s operating platform will be unavailable for more than 48 hours;
    • The cloud environment hosting conductor or Foundry’s FHIR service is damaged and will be unavailable for more than 24 hours;
    • Other criteria as appropriate and as defined by Foundry leadership at the time.
  • If the plan is activated, the Security Officer notifies team members and informs them of the details of the event and whether relocation is required.
  • Group leaders notify their respective teams. Team members are informed of all applicable information and prepared to respond and relocate if necessary.
  • The Security Officer notifies vendor and infrastructure partners (AWS support and any subprocessors required for recovery) that a contingency event has been declared.
  • The Security Officer notifies remaining personnel and executive leadership of the general status of the incident.
  • Notification can be made by Slack, email, or phone, in that order of preference.
  • If PHI is or may be implicated, the Breach Notification workflow is invoked in parallel, notifying affected practices per the BAA.

8.2 Recovery phase

These procedures cover recovering Foundry infrastructure and operations at an alternate site (typically an alternate AWS region), while other efforts are directed to repair damage to the original system and capabilities. Each procedure should be executed in the sequence presented, with parallel work permitted where indicated.

Recovery goal: rebuild Foundry’s operating platform to a production state, with PHI restored from the latest authorized FHIR export.

  1. Contact affected practices to begin initial communication (Platform Engineering, with Communications Lead).
  2. Assess damage to the environment (Platform Engineering).
  3. Create a new production environment using infrastructure-as-code (Terraform/Pulumi) and the new-environment bootstrap automation (Platform Engineering).
  4. Ensure secure access to the new environment, including identity federation, secrets distribution, and break-glass account provisioning (Security).
  5. Begin code deployment and data restoration:
    • Application services and SSR web tier from container registry / source (Platform Engineering);
    • FHIR R4 PHI restoration via Foundry’s FHIR service import from the most recent encrypted export (Platform Engineering, with Security oversight).
  6. Test the new environment and applications using pre-written tests, including smoke tests, schema validation, and PHI sampling validation (without exporting PHI to test outputs) (Platform Engineering).
  7. Test logging, security alerting, and audit logging functionality (Platform Engineering and Security).
  8. Confirm systems and applications are patched and up to date (Platform Engineering).
  9. Update DNS and other necessary records to point to the new environment (Platform Engineering).
  10. Update affected practices through established channels and the Foundry status page (Communications Lead).

For detailed recovery instructions, consult the Backup and Recovery section of Data Management and the platform-internal runbook (link held in the Engineering wiki).

8.3 Reconstitution phase

This section discusses activities necessary for restoring full Foundry operations at the original or new region. The goal is to provide a controlled transition of operations from the alternate site back to the primary site (or to formally adopt the alternate site as the new primary).

  1. Original or new site restoration. Repeat the relevant Recovery-phase steps at the chosen primary site. Restoration of the original site is unnecessary for purely cloud-hosted environments except when required for forensic purposes.
  2. Plan deactivation. If the Foundry environment is moved back to the original site from the alternate site, any temporary cloud resources used during the alternate-site operation are decommissioned per Foundry’s media disposal practices, including verified destruction of any storage used to stage PHI exports during recovery.

9. Production environments and PHI recovery

PHI lives only in Foundry’s FHIR service (R4) store. Backup, replication, and recovery for the PHI store follow Foundry’s FHIR service procedures:

  • Backup posture. the FHIR service store use our provider’s intra-region durability; managed object storage exports run on a continuous schedule into a dual-region bucket with object versioning and a retention policy aligned with HIPAA recordkeeping. The recovery posture is a fast in-region rebuild from infrastructure-as-code plus a managed object storage replay; the dual-region export protects against any single-region loss.
  • Replication. Where supported by AWS, exports are replicated to an alternate region distinct from the primary production region.
  • Encryption. Exports are encrypted at rest with AES-256, using provider-managed or customer-managed encryption keys (CMEK) when CMEK is selected by the practice via BAA addendum.
  • Restoration. In a worst-case scenario, complete loss of the production project, Foundry rebuilds the project from infrastructure-as-code and restores PHI from the most recent authorized FHIR export held in the alternate-region bucket. The restored store is validated against pre-defined integrity checks before being placed back into service.
  • No PHI on local disk. Workstations, agent worktrees, and Foundry-owned tooling never store PHI. Recovery does not depend on, and does not produce, any local PHI artifact.

Detailed recovery procedures are in the Backup and Recovery section of Data Management and the Engineering wiki.

10. 90-day clean offboarding

Independent of disaster events, Foundry commits to a clean 90-day offboarding for any practice that chooses to discontinue services. Offboarding is treated as a planned continuity event:

  • The practice notifies Foundry of its intent to offboard. Foundry confirms in writing and assigns an offboarding lead.
  • Within the 90-day window, Foundry produces a complete export of the practice’s PHI from the FHIR store, in a portable format (FHIR R4 bundle), and delivers it through the channel agreed in the BAA.
  • Foundry completes any operational handoffs required for the practice to resume independent operation, including documentation and identity offboarding.
  • At the end of the 90-day window, Foundry returns or destroys all PHI in its possession in accordance with the BAA and HIPAA § 164.504(e)(2)(ii)(J). Destruction is documented with a certificate of destruction; return is documented with confirmation of receipt from the practice.
  • Foundry retains documentation of the offboarding event for the HIPAA retention period.

11. Work-site recovery

If a Foundry facility is not functioning due to a disaster, employees work from home or from a secondary site with Internet access until the physical recovery of the affected facility is complete. Recovery is performed by the facility-management firm under contract with Foundry and coordinated by the Facility Manager or site lead.

Foundry’s engineering organization is location-independent and does not require an office-provided Internet connection to continue operating. Workstations are MDM-managed and operate over the public internet to Foundry-controlled services; no on-prem network access is required for normal work.

12. Service event communications

Foundry maintains a status page that provides real-time updates on service status. The status page is updated with details about events that may cause service interruption or downtime. A follow-up root-cause analysis (RCA) is made available to affected practices upon request after the event has been resolved.

Short event (hours)

  • Practices experience a short delay in service.
  • Foundry monitors the event and determines course of action. Escalation may be required.

Moderate event (days)

  • Practices experience a modest delay in service; processes running may need to be restarted.
  • Foundry monitors the event and determines course of action. Escalation may be required.
  • Foundry notifies affected practices of the delay and provides updates on the status page.

Long event (a week or more)

  • Practices experience an extended delay in service; processes running may need to be restarted.
  • Foundry monitors the event and determines course of action. Escalation may be required.
  • Foundry notifies affected practices of the delay and provides regular updates on the status page; if PHI is implicated, the Breach Notification workflow applies.

13. Testing & maintenance

The Security Officer and the Engineering Lead establish criteria for validation and testing of this Contingency Plan, set an annual test schedule, and ensure implementation of the test. The testing process also serves as training for personnel involved in execution. At a minimum, the Contingency Plan is tested annually (within 365 days).

Validation and testing exercises include tabletop and technical testing. All application systems are covered by at least the tabletop process. If an application system’s contingency plan is included in the technical testing of a supporting system, that technical test satisfies the annual requirement for the application.

13.1 Tabletop testing

Tabletop testing validates that designated personnel are knowledgeable and capable of performing the notification/activation requirements and procedures in a timely manner. Exercises include, but are not limited to:

  • Simulating a specific crisis to test the ability to respond in a coordinated, timely, and effective manner;
  • Walking through the line of succession with key roles temporarily unavailable;
  • Scenario play covering an agent runaway, an Atlas control-plane compromise, and a regional AWS outage.

13.2 Technical testing

Technical testing ensures that communication processes and data storage and recovery processes can function at an alternate site to perform the functions and capabilities of the system within the designated requirements. Technical testing includes, but is not limited to:

  • Processing from a backup system at the alternate site;
  • Restoring the FHIR store from the most recent encrypted export into an alternate-region store;
  • Switching compute and storage resources to the alternate processing site.

Findings from each exercise are tracked to closure in Linear and feed into the next annual revision of the plan.

14. Roles & responsibilities

  • Security Officer: authority for the entire Contingency Plan; activates the plan; coordinates with executive leadership; chairs post-event review.
  • Engineering Lead: leads technical recovery; owns the rebuild-from-code procedures and FHIR restoration; co-authority for activation.
  • Platform Engineering: executes recovery procedures.
  • Security: assesses cyber incidents per the IR policy; secures alternate sites during recovery.
  • People & Facilities: physical safety, site recovery, work-site logistics.
  • Communications Lead (per-event role): updates affected practices and the status page.
  • CEO: line of succession; external communications above a defined threshold.
  • Workforce members: follow the activation instructions issued by their team lead.

15. Review & revision

This policy and the underlying Contingency Plan are reviewed at least annually, after every executed activation, and whenever a significant infrastructure or organizational change makes the plan stale. Material revisions are approved by the Security Officer in coordination with the Engineering Lead and the Policy Management process.

16. Related policies