This repository has been archived on 2026-07-29. You can view files and clone it, but you cannot make any changes to it's state, such as pushing and creating new issues, pull requests or comments.
astro-pipeline/README.md
laurence 0b6c42e4c8 Point this repository at its new home
Merged into git.discworld.casa/laurence/astrophotography on 2026-07-21,
where it lives on as the pipeline/ directory alongside the telescope
network reference, the drain campaign and the observing plans, with both
histories preserved.

This repository was three hours old at the point of merging, which is
itself the argument for merging: it was created to get the processing
code out of the image directory, not because the code was a separate
concern from the imaging it processes.
2026-07-21 17:17:37 +01:00

110 lines
4.8 KiB
Markdown

> [!IMPORTANT]
> **This repository has moved.** Since 2026-07-21 it lives on as the
> `pipeline/` directory of
> **[astrophotography](https://git.discworld.casa/laurence/astrophotography)**,
> merged with the iTelescope network reference and campaign. Both histories
> were preserved.
>
> Do new work there, not here.
# astro-pipeline
Processing and analysis code for remote-telescope imaging sessions, starting
with iTelescope data from the [itelescope](https://git.discworld.casa/laurence/itelescope)
drain campaign.
The code lives here. **The data does not** - image sessions stay on disk (or
wherever they are archived) and are addressed by an environment variable, so a
session directory contains only pixels, results and a description of what was
done to them.
## What is here now
`session-scripts/` - the 50 scripts that processed the NGC 5128 session of
2026-07-21, exactly as they were run, plus the shared `layout.py` that tells
them where files live. This is a working record rather than a finished product:
the scripts were written in sequence as the work went along, several of them by
parallel agents, and they show it. They are kept because they are the honest
provenance of a set of published results, and because the productionised
pipeline should be able to reproduce those results exactly.
`session-scripts/restructure.py` - reorganises a flat session directory into the
named layout below. Idempotent, dry run by default.
`observing/` - plans for observing sessions that are not remote-telescope runs.
Currently the total solar eclipse of 12 August 2026, seen from Menorca. These
live here rather than in a notes app because they are worked out from real
numbers, they get revised as the date approaches, and the reasoning behind each
decision is worth keeping.
## Pointing the scripts at a session
```
set ASTRO_SESSION=D:\astro\NGC5128\20260721 # Windows
export ASTRO_SESSION=/data/astro/NGC5128/20260721 # POSIX
python session-scripts/layout.py # prints the resolved layout
```
`layout.py` maps a **filename** to its subdirectory, so a script asks for
`master-Red.fit` or `_stars.npz` and gets the right path without knowing the
directory structure:
| Directory | Holds |
|---|---|
| `raw/` | exactly what the telescope delivered: archives and their preview jpegs |
| `calibrated/` | uncompressed calibrated subs |
| `stacks/masters/` | per-filter registered, plate-solved masters |
| `stacks/original/` | alignment-only baseline stacks, no other processing |
| `final/` | the deliverable renderings |
| `renderings/` | other finished images |
| `science/figures/` | analysis plots |
| `science/catalogues/` | measured tables (CSV) |
| `science/data/` | models, masks, derived quantities |
| `science/notes/` | analysis write-ups |
| `intermediates/` | caches a re-run can regenerate |
Every session directory also carries its own `METHODS.md` describing what was
done to that data and what was found - written for a reader who was not there.
## Running order
The scripts are named for their stage and run in this order:
```
unzip.py -> analyse.py -> stack.py -> solve.py -> depth.py
-> compose.py -> hdr.py / enhance.py / starless.py / annotate.py
-> final.py -> closeup.py -> triptych.py
```
The analysis families are independent of each other and of the renderings:
`gc-*` (globular clusters), `sb_*` (surface photometry), `mo_*` (moving objects
and transients).
## Requirements
Python 3.12 with numpy, scipy, astropy, scikit-image, sep, astroalign,
photutils, astroquery, matplotlib, tifffile, Pillow.
## Where this is going
The next piece of work is a scheduler-driven pipeline: a staged CLI
(`ingest -> calibrate -> measure -> register -> stack -> solve -> compose ->
analyse`) with each stage resumable, packaged as an Apptainer image and driven
by Slurm array jobs. Targets beyond mono LRGB: narrowband palettes, one-shot
colour with debayering, other observatories' header conventions, and full
calibration from bias/dark/flat for sources that do not pre-calibrate.
Three findings from the first session are requirements for that build, not
optional extras:
1. **Vet moving-object candidates in detector coordinates.** Registration holds
the sky still, so it drags detector-fixed defects across the frame on
perfectly straight, constant-rate tracks. Hot pixels are better-behaved
asteroids than real asteroids. This one cut took 141 confident spurious
detections to zero.
2. **Carry `r50/psf` through to any catalogue cross-match.** Comparing an
aperture magnitude of a resolved source against a point-source catalogue
like Gaia is meaningless, and looks exactly like a 2.8 magnitude outburst.
3. **Never fit a sky background to a field the target fills.** A plane fitted
around a large galaxy absorbs its halo - measured at -17.9 ADU/px here.
Fit the background and a source model together.