# NULISA pilot: bridge and control wells, raw reports (4 plates)

Raw data behind the variance-component calibration (Table 1) of

> Abe K, Nguyen TT, Otiti I, Kariuki S, Maecker HT. *How many bridge specimens are
> enough? A practical cost-vs-precision framework for multi-batch high-plex
> proteomics.* Journal of Proteome Research (submitted).

Analysis code: doi:10.5281/zenodo.21895289 (specimen-bridging-nulisa, v2.0 or later).

## What is in this deposit

Four NULISAseq Inflammation Panel runs (250 protein targets, one 96-well plate per run). From each plate, only the 16 wells that carry no study-participant
material are included:

| Wells per plate | Label | What it is |
|---|---|---|
| 6 | `BRIDGE-Pilot p-k` (NPQ sheet), `BRIDGE-Pp-k` (other sheets) | The bridge specimen: one control serum run in six replicate wells on every plate. `p` = plate 1 to 4, `k` = replicate 1 to 6 |
| 4 | `NC_1` to `NC_4` | Negative controls |
| 3 | `SC_1` to `SC_3` | Sample controls |
| 3 | `IPC_1` to `IPC_3` | Inter-plate controls |

64 wells in total. The other 80 wells on each plate held study samples and are **not**
included (see "What was removed").

## Files

| File | Contents |
|---|---|
| `09-10-24 NULISAseq_Inflammation_Pilot Run 1_Report V1.xlsx` | Plate 1 report |
| `09-10-24 NULISAseq_Inflammation_Pilot Run 2_Report V1.xlsx` | Plate 2 report |
| `09-13-24 NULISAseq_inflammation_Pilot Run 3_Report V1.xlsx` | Plate 3 report |
| `09-13-24 NULISAseq_Inflammation_Pilot Run 4_Report V1.xlsx` | Plate 4 report |
| `bridge_and_control_NPQ_long.csv` | The NPQ values of all 64 wells in one long table |
| `MANIFEST.sha256` | SHA-256 checksums of the files above |

Each workbook keeps the four sheets of the vendor report, in the original order:

- **NPQ**: NULISA Protein Quantification (normalized, log2 scale). One row per target, one column per well.
- **Raw-counts**: sequencing read counts. One row per target plus the `mCherry` internal control, one column per well.
- **Detectability**: per-target detectability and limit of detection (`targetLOD` on the count scale, `targetLOD_NPQ` on the NPQ scale).
- **Sample and Run QC**: metric definitions, then one row per retained well (IC median, detectability, IC reads, QC status, plate row and column, reads), and a run-level summary (Run ID, Run QC, wells passed, IC/IPC CVs, run detectability, total reads).

The workbook file names and well labels match what the analysis code expects, so the
files can be used as they are.

`bridge_and_control_NPQ_long.csv` columns: `batch` (Run1 to Run4), `well`, `target`,
`value_npq_log2`, `well_type` (`bridge` or `control`).

## What was removed, and why

The 320 study samples come from a clinical study whose results are not yet published.
Their wells were removed from every sheet, and study identifiers were replaced with neutral
labels: the bridge specimen is labelled `BRIDGE`, and the study name was taken out of the
Run IDs and file names. Values were not changed.

Two items in the reports are **plate-level summaries computed by the vendor software over
all 96 wells**, study samples included: the Detectability sheet, and the run-level block of
the Sample and Run QC sheet. They are kept because they describe the plate as a whole and
contain no per-sample values.

**None of the paper's reported results depend on the removed wells.** Table 1 is estimated
from the bridge wells alone. The design simulation draws baseline protein profiles from
study samples, but every result it reports is a difference between plates, and the same
baseline is added to both plates, so it cancels.

## Reproducing Table 1

```sh
# after downloading the code deposit (doi:10.5281/zenodo.21895289) and unzipping it
cp *.xlsx specimen-bridging-nulisa-v2.0/data/pilot_runs/
cd specimen-bridging-nulisa-v2.0/R
Rscript 1_calibration_nulisa_revised.R
```

This regenerates `results/week1_revised/variance_components_by_target.csv`. Checked
against the published file on 24 September 2026: all 250 targets agree to within 1e-14
(relative), with `batch_var` identical.

## License

CC BY 4.0.
