diff --git a/CLAUDE.md b/CLAUDE.md index 2b0c8043..713a5dc6 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -79,7 +79,7 @@ problem — read the `--check` diff and apply it by hand rather than fighting th | `data-exchange` | spreadsheet and Google Sheets imports, desktop Google Sheets integration, preview | | `file-ingest` | filesystem inventory and duplicate resolution for folder-backed projects and live imports. Only lists paths and decides duplicates; never maps columns or touches a project | | `updater` | signed update manifests, install provenance, and the `qrate-update-helper` binary that applies an update after qrate exits | -| `qrate-export` | shared project reader and CSV/Excel/JSON-LD/CSL-JSON/ZIP writers; optional browser WASM API | +| `qrate-export` | shared project reader and CSV/Excel/JSON-LD/CSL-JSON/IIIF/ZIP writers; optional browser WASM API | | `diagnostics` | the validator, spelling checks, fixes, and the problems panel | | `checks` | date and authority validators, registered by `app` through the `diagnostics` crate | | `spellcheck` | dictionary catalogue behind `diagnostics::spelling` | diff --git a/README.md b/README.md index 17262510..ee82bbaf 100644 --- a/README.md +++ b/README.md @@ -24,7 +24,7 @@ Collection data rarely lives in just one place. A spreadsheet names an object, a - **See the records and the material together.** Link by filename or a pattern, browse a gallery of thumbnails, and preview images, documents, audio, and video alongside each record. - **Find pictures by what they show.** Optional visual search ranks linked images, PDF pages, and video frames against a plain description, or against another record's picture. The model downloads once and runs on your own machine. - **Catch problems while you work.** Check spelling, date formats, file links, headings, and selected authority sources. Review a proposed whole-cell correction before applying it. -- **Take your data where it needs to go.** Export CSV, Excel, JSON-LD, CSL-JSON, or a ZIP archive; Excel keeps your notes as cell comments. Google Sheets export and sync are available when you choose to enable them. +- **Take your data where it needs to go.** Export CSV, Excel, JSON-LD, CSL-JSON, a IIIF manifest, or a ZIP archive; Excel keeps your notes as cell comments. Google Sheets export and sync are available when you choose to enable them. - **Use AI with a human in control.** An optional local agent can review the open project and stage findings in the Problems panel. It cannot change a cell; you decide what to accept. ## A typical workflow diff --git a/crates/ai/src/traits.rs b/crates/ai/src/traits.rs index 5f0d545e..a2edab0c 100644 --- a/crates/ai/src/traits.rs +++ b/crates/ai/src/traits.rs @@ -50,6 +50,8 @@ pub struct ReviewerCapabilities { /// What the agent panel asks of a model: does this row describe this image, and where doesn't it. /// /// `row_data` is the row as JSON — column name to value — so a provider needs no notion of grids. +// `async_trait` writes the `#[must_use]` that clippy 1.99 objects to. +#[allow(clippy::double_must_use)] #[async_trait] pub trait DataReviewer: Send + Sync { async fn initialize(&self) -> Result<()>; diff --git a/crates/app/src/export.rs b/crates/app/src/export.rs index b28ce414..e2748f62 100644 --- a/crates/app/src/export.rs +++ b/crates/app/src/export.rs @@ -2,7 +2,7 @@ //! //! Everything the writers need is read out of `cx` before any dialog opens, so the spawned task //! only carries plain values. The formats themselves live in `qrate-export`; what's here is the -//! action, the save dialog, and the CSL field-mapping picker. +//! action, the save dialog, the CSL field-mapping picker, and the IIIF address prompt. use std::cell::RefCell; use std::fs::File; @@ -19,11 +19,12 @@ use gpui::{ }; use gpui_component::button::Button; use gpui_component::dialog::DialogButtonProps; +use gpui_component::input::{Input, InputState}; use gpui_component::menu::{DropdownMenu as _, PopupMenu, PopupMenuItem}; use gpui_component::notification::Notification; use gpui_component::{Sizable as _, WindowExt as _, h_flex}; use qrate_export::export::{self, ArchiveFile, CSL_FIELDS, CslMapping, ExportComponent}; -use qrate_export::{ProjectNote, SheetNote}; +use qrate_export::{IiifInput, ProjectNote, SheetNote}; use schemars::JsonSchema; use serde::Deserialize; use settings::columns::ColumnType; @@ -46,6 +47,9 @@ pub const CSV_DELIMITERS: &[(&str, &str)] = &[ ("tab", "Tab"), ]; +/// `__settings` key for the web address the open project's IIIF manifest is published at. +const IIIF_BASE_KEY: &str = "iiif_base_url"; + /// `__settings` key for the folder the open project last exported into. const EXPORT_FOLDER_KEY: &str = "last_export_folder"; @@ -93,8 +97,10 @@ struct ExportGrid { /// Each column's note, for the formats that carry notes; the row notes are read from the /// project file off the UI thread. column_notes: Vec<(String, String)>, - /// Where a ZIP finds each row's linked file. `None` for every other format. + /// Where a ZIP or a IIIF manifest finds each row's linked file. `None` for the other formats. files: Option, + /// What a IIIF manifest is built on. `None` for every other format. + iiif: Option, csv: export::CsvOptions, } @@ -104,12 +110,20 @@ struct ZipFiles { declared: Vec, } +/// The web address a manifest is published at, the project's name, and each column's type. +struct IiifTarget { + base_url: String, + label: String, + columns: Vec<(String, ColumnType)>, +} + #[derive(Clone, Copy, PartialEq, Eq, Deserialize, JsonSchema)] pub enum ExportFormat { Csv, Xlsx, JsonLd, Csl, + Iiif, Zip, GoogleSheet, GoogleSheetSync, @@ -134,11 +148,12 @@ const MAX_PLUGIN_EXPORT_BYTES: usize = 64 * 1024 * 1024; /// Menu order. Each entry is the label and the extension of the file the save dialog offers, which /// is named after the project; the Sheets target never touches disk, so it has no name to suggest. -pub const EXPORT_FORMATS: [(ExportFormat, &str, Option<&str>); 7] = [ +pub const EXPORT_FORMATS: [(ExportFormat, &str, Option<&str>); 8] = [ (ExportFormat::Csv, "CSV…", Some("csv")), (ExportFormat::Xlsx, "Excel (.xlsx)…", Some("xlsx")), (ExportFormat::JsonLd, "JSON-LD…", Some("jsonld")), (ExportFormat::Csl, "Zotero (CSL-JSON)…", Some("json")), + (ExportFormat::Iiif, "IIIF Manifest…", Some("json")), (ExportFormat::Zip, "ZIP Archive…", Some("zip")), (ExportFormat::GoogleSheet, "New Google Sheet…", None), (ExportFormat::GoogleSheetSync, "Sync to Google Sheet…", None), @@ -342,7 +357,7 @@ pub fn run(format: ExportFormat, window: &mut Window, cx: &mut App) { .iter() .map(|column| (column.name.clone(), column.notes.clone())) .collect::>(); - let files = (format == ExportFormat::Zip).then(|| ZipFiles { + let files = matches!(format, ExportFormat::Zip | ExportFormat::Iiif).then(|| ZipFiles { folder: project .data .values @@ -351,6 +366,16 @@ pub fn run(format: ExportFormat, window: &mut Window, cx: &mut App) { .unwrap_or_default(), declared: table::photos::declared_file_columns(&project.data), }); + let iiif = (format == ExportFormat::Iiif).then(|| IiifTarget { + base_url: project + .data + .values + .get(IIIF_BASE_KEY) + .map(|v| v.text().to_string()) + .unwrap_or_default(), + label: title.clone(), + columns: declared, + }); let grid = ExportGrid { headers, row_ids, @@ -358,6 +383,7 @@ pub fn run(format: ExportFormat, window: &mut Window, cx: &mut App) { structure, column_notes, files, + iiif, csv: csv_options(cx), }; // A project that already knows its spreadsheet refills that one; otherwise "Sync" asks @@ -383,6 +409,9 @@ pub fn run(format: ExportFormat, window: &mut Window, cx: &mut App) { if format == ExportFormat::Csl { return ask_csl_mapping(file, title, grid.headers, grid.rows, window, cx); } + if format == ExportFormat::Iiif { + return ask_iiif_address(file, title, grid, window, cx); + } save_as(format, file, title, grid, CslMapping::new(), cx); } @@ -410,17 +439,18 @@ fn sheet_notes(file: &Path, grid: &ExportGrid) -> anyhow::Result> /// Each row's linked file, and where it sat in the source tree. Scans the files folder, so it runs /// off the UI thread. -fn archive_files(grid: &ExportGrid, files: &ZipFiles) -> Vec { +fn archive_files(grid: &ExportGrid, files: &ZipFiles) -> Vec> { let root = Path::new(&files.folder); table::photos::resolve_row_images(&grid.headers, &grid.rows, &files.folder, &files.declared) .into_iter() - .flatten() - .map(|path| ArchiveFile { - source_path: Some(match qrate_export::relative_to(root, &path) { - Some(relative) => relative.to_string_lossy().replace('\\', "/"), - None => path.to_string_lossy().into_owned(), - }), - path, + .map(|path| { + path.map(|path| ArchiveFile { + source_path: Some(match qrate_export::relative_to(root, &path) { + Some(relative) => relative.to_string_lossy().replace('\\', "/"), + None => path.to_string_lossy().into_owned(), + }), + path, + }) }) .collect() } @@ -513,11 +543,15 @@ fn save_as( ) { let window = cx.active_window(); let directory = export_folder(&project_file, cx); - let suggested = EXPORT_FORMATS - .iter() - .find(|(f, _, _)| *f == format) - .and_then(|(_, _, extension)| *extension) - .map(|extension| suggested_name(&title, extension)); + let suggested = match format { + // The manifest's own address ends in this name. + ExportFormat::Iiif => Some("manifest.json".to_string()), + _ => EXPORT_FORMATS + .iter() + .find(|(f, _, _)| *f == format) + .and_then(|(_, _, extension)| *extension) + .map(|extension| suggested_name(&title, extension)), + }; let receiver = cx.prompt_for_new_path(&directory, suggested.as_deref()); cx.spawn(async move |cx| { @@ -530,13 +564,13 @@ fn save_as( let task = cx.background_spawn({ let (path, progress) = (path.clone(), progress.clone()); async move { - let result = (|| -> anyhow::Result<()> { + let result = (|| -> anyhow::Result> { let parent = path.parent().unwrap_or_else(|| Path::new(".")); let temporary = tempfile::Builder::new() .prefix(".qrate-export-") .tempfile_in(parent)? .into_temp_path(); - write_export( + let outcome = write_export( format, &project_file, &temporary, @@ -548,7 +582,7 @@ fn save_as( anyhow::bail!(CANCELLED); } temporary.persist(&path)?; - Ok(()) + Ok(outcome) })(); progress.done.store(true, Ordering::Release); result @@ -587,7 +621,8 @@ fn save_as( |name| name.to_string_lossy().into_owned(), ); let note = match task.await { - Ok(()) => Notification::success(format!( + Ok(Some(outcome)) => Notification::success(outcome), + Ok(None) => Notification::success(format!( "Exported {rows} row{} to {name}", if rows == 1 { "" } else { "s" } )), @@ -604,7 +639,8 @@ fn save_as( .detach(); } -/// Write `grid` to `path` in `format`. Runs on the background executor. +/// Write `grid` to `path` in `format`, answering with what to tell the archivist when "exported +/// every row" is not the whole story. Runs on the background executor. fn write_export( format: ExportFormat, project_file: &Path, @@ -612,7 +648,7 @@ fn write_export( grid: &ExportGrid, mapping: &CslMapping, progress: &Arc, -) -> anyhow::Result<()> { +) -> anyhow::Result> { let ExportGrid { headers, row_ids, @@ -630,11 +666,35 @@ fn write_export( &export::jsonld_hierarchy_value(headers, row_ids, rows, structure), )?, ExportFormat::Csl => export::write_json(path, &export::csl_items(headers, rows, mapping))?, + ExportFormat::Iiif => { + let (Some(target), Some(files)) = (&grid.iiif, &grid.files) else { + anyhow::bail!("the IIIF export was started without its web address"); + }; + let (manifest, issues) = qrate_export::iiif_manifest(&IiifInput { + base_url: &target.base_url, + label: &target.label, + headers, + row_ids, + rows, + structure, + columns: &target.columns, + media: &qrate_export::iiif_media(&archive_files(grid, files), preview::extent), + })?; + export::write_json(path, &manifest)?; + let shown = manifest["items"].as_array().map_or(0, Vec::len); + let mut outcome = format!("Exported {shown} of {} rows to a IIIF manifest", rows.len()); + let summary = qrate_export::iiif_summary(&issues); + if !summary.is_empty() { + outcome = format!("{outcome}: {summary}"); + } + log::info!("{outcome}"); + return Ok(Some(outcome)); + } ExportFormat::Zip => { - let images = grid + let images: Vec = grid .files .as_ref() - .map(|files| archive_files(grid, files)) + .map(|files| archive_files(grid, files).into_iter().flatten().collect()) .unwrap_or_default(); let total: u64 = images .iter() @@ -651,7 +711,77 @@ fn write_export( // Handled in `run` — they have no path to write to. ExportFormat::GoogleSheet | ExportFormat::GoogleSheetSync => {} } - Ok(()) + Ok(None) +} + +/// Where the manifest will be published. Every address in a IIIF manifest is a web address and +/// qrate hosts nothing, so only the archivist knows it. Opens on the project's last answer. +fn ask_iiif_address( + project_file: PathBuf, + title: String, + grid: ExportGrid, + window: &mut Window, + cx: &mut App, +) { + let saved = grid + .iiif + .as_ref() + .map(|target| target.base_url.clone()) + .unwrap_or_default(); + let input = cx.new(|cx| { + InputState::new(window, cx) + .placeholder("https://example.org/iiif/my-collection") + .default_value(saved) + }); + // The dialog's builder re-runs on every frame, but the grid is handed over once. + let grid = Rc::new(RefCell::new(Some(grid))); + window.open_dialog(cx, move |dialog, _, _| { + let (input, for_ok) = (input.clone(), input.clone()); + let (grid, project_file, title) = (grid.clone(), project_file.clone(), title.clone()); + dialog + .title("Export a IIIF Manifest") + .w(gpui::px(460.0)) + .content(move |content, _, _| { + content + .p_4() + .gap_2() + .child( + "The web address of the folder this manifest will be published in. Put \ + the linked files in a files folder beside it, laid out as the ZIP \ + Archive export lays them out.", + ) + .child(Input::new(&input).w_full()) + }) + .button_props(DialogButtonProps::default().ok_text("Export…")) + .on_ok(move |_: &ClickEvent, window, cx| { + let base_url = for_ok.read(cx).value().trim().to_string(); + if let Err(err) = qrate_export::iiif::base_url(&base_url) { + window.push_notification( + Notification::error(err.to_string()).id::(), + cx, + ); + return false; + } + let Some(mut grid) = grid.borrow_mut().take() else { + return true; + }; + if cx.has_global::() { + CurrentProject::set_text(IIIF_BASE_KEY, base_url.clone().into(), cx); + } + if let Some(target) = grid.iiif.as_mut() { + target.base_url = base_url; + } + save_as( + ExportFormat::Iiif, + project_file.clone(), + title.clone(), + grid, + CslMapping::new(), + cx, + ); + true + }) + }); } /// Which column feeds which CSL field. Opens on the saved answer, or on what the declared column @@ -768,6 +898,7 @@ fn ask_csl_mapping( structure: Vec::new(), column_notes: Vec::new(), files: None, + iiif: None, csv: export::CsvOptions::default(), }, mapping, @@ -1097,3 +1228,74 @@ async fn sign_in( } } } + +#[cfg(test)] +mod tests { + use std::path::Path; + use std::sync::Arc; + + use qrate_export::export::CslMapping; + use settings::columns::ColumnType; + + use super::{ExportFormat, ExportGrid, IiifTarget, Progress, ZipFiles, write_export}; + + /// The sample collection's scans are real JPEGs, so this measures files the way an export does. + #[test] + fn a_iiif_export_measures_the_linked_files_and_reports_the_rest() { + let photos = Path::new(env!("CARGO_MANIFEST_DIR")).join("../../sample/photos"); + let row = |id: &str, caption: &str| vec![id.to_string(), caption.to_string()]; + let grid = ExportGrid { + headers: vec!["Digital ID".into(), "Caption".into()], + row_ids: vec![1, 2, 3], + rows: vec![ + row("1", "Yad Vashem Memorial"), + row("2", "Dome of the Rock"), + row("999", "Never scanned"), + ], + structure: Vec::new(), + column_notes: Vec::new(), + files: Some(ZipFiles { + folder: photos.to_string_lossy().into_owned(), + declared: vec!["Digital ID".into()], + }), + iiif: Some(IiifTarget { + base_url: "https://example.org/iiif/aderman".into(), + label: "Aderman".into(), + columns: vec![ + ("Digital ID".into(), ColumnType::Filename), + ("Caption".into(), ColumnType::Title), + ], + }), + csv: Default::default(), + }; + let out = tempfile::tempdir().unwrap(); + let path = out.path().join("manifest.json"); + let outcome = write_export( + ExportFormat::Iiif, + Path::new("unused.qrate"), + &path, + &grid, + &CslMapping::new(), + &Arc::new(Progress::default()), + ) + .unwrap(); + assert_eq!( + outcome.as_deref(), + Some("Exported 2 of 3 rows to a IIIF manifest: left out 1 row with no linked file") + ); + + let manifest: serde_json::Value = + serde_json::from_slice(&std::fs::read(&path).unwrap()).unwrap(); + let canvases = manifest["items"].as_array().unwrap(); + assert_eq!(canvases.len(), 2); + let (width, height) = preview::dimensions(&photos.join("1.jpg")).unwrap(); + assert!(width > 0 && height > 0); + assert_eq!(canvases[0]["width"], width); + assert_eq!(canvases[0]["height"], height); + assert_eq!(canvases[0]["label"]["none"][0], "Yad Vashem Memorial"); + assert_eq!( + canvases[0]["items"][0]["items"][0]["body"]["id"], + "https://example.org/iiif/aderman/files/1.jpg" + ); + } +} diff --git a/crates/preview/src/lib.rs b/crates/preview/src/lib.rs index fb00b3c4..10aab9b5 100644 --- a/crates/preview/src/lib.rs +++ b/crates/preview/src/lib.rs @@ -358,6 +358,28 @@ pub fn dimensions(path: &Path) -> Option<(u32, u32)> { }) } +/// How much room a file takes when it is shown, for whichever it has: a picture's pixels, a +/// document's first page in points, a recording's seconds, a video's both. +/// +/// Opens the file, and spawns ffmpeg for a video, so ask it off the UI thread. +pub fn extent(path: &Path) -> (Option<(u32, u32)>, Option) { + let Some(extension) = extension(path) else { + return (None, None); + }; + if pdf::handles(&extension) { + return (pdf::page_size(path), None); + } + if audio::handles(&extension) { + let length = audio::duration(path).map(|length| length.as_secs_f64()); + return (None, length); + } + if media::handles(&extension) { + let (seconds, size) = media::probe(path); + return (size, seconds.filter(|_| media::is_video(&extension))); + } + (dimensions(path), None) +} + /// `2.4 MB`. Powers of 1024 with the unit names every file manager on the three platforms shows, /// and whole bytes below a kilobyte — "0.3 KB" reads as a rounding of something, not as a stub. pub fn file_size(bytes: u64) -> String { @@ -412,7 +434,10 @@ pub fn has_video(path: &Path) -> bool { /// /// Spawns ffmpeg, so it costs about a tenth of a second: ask it when a viewer opens, never per row. pub fn video_duration(path: &Path) -> Option { - has_video(path).then(|| media::duration(path)).flatten() + has_video(path) + .then(|| media::probe(path).0) + .flatten() + .map(|seconds| seconds as u32) } /// Whether a recording carries artwork, and so whether there is anything to look at while it @@ -1411,6 +1436,32 @@ mod tests { }); } + /// A picture answers with pixels and a recording with seconds, never the other's. + #[test] + fn a_file_reports_the_extent_its_kind_has() { + use image::ImageEncoder as _; + + let picture = std::env::temp_dir().join("qrate-extent-probe.png"); + image::codecs::png::PngEncoder::new(std::fs::File::create(&picture).unwrap()) + .write_image(&[0u8; 4 * 3 * 2], 3, 2, image::ExtendedColorType::Rgba8) + .unwrap(); + assert_eq!(crate::extent(&picture), (Some((3, 2)), None)); + + // Two seconds of 8 kHz silence. + let recording = std::env::temp_dir().join("qrate-extent-probe.wav"); + std::fs::write(&recording, crate::playback::silent_wav(16000)).unwrap(); + let (size, seconds) = crate::extent(&recording); + assert_eq!(size, None); + assert!((seconds.expect("a valid WAV has a length") - 2.0).abs() < 0.01); + + assert_eq!( + crate::extent(Path::new("/nonexistent/notes.docx")), + (None, None) + ); + let _ = std::fs::remove_file(&picture); + let _ = std::fs::remove_file(&recording); + } + /// A phone photo stored sideways with an EXIF turn has to come out upright in every thumbnail, /// as it does in the viewer, where gpui applies the tag itself. The viewer's own rotation is /// applied on top of that, never instead of it. diff --git a/crates/preview/src/media.rs b/crates/preview/src/media.rs index 362112f8..68aed80b 100644 --- a/crates/preview/src/media.rs +++ b/crates/preview/src/media.rs @@ -130,32 +130,53 @@ pub fn frame_at(path: &Path, max_edge: u32, seconds: u32) -> Option Option { - let output = Command::new(binary()?) - .arg("-nostdin") - .arg("-i") - .arg(path) - .stdin(Stdio::null()) - .stdout(Stdio::null()) - .stderr(Stdio::piped()) - .output() - .ok()?; - - // `Duration: 00:04:03.20, start: ...`. Unparseable for a stream with no known length, which - // reads as "no timeline" rather than as an error. +pub fn probe(path: &Path) -> (Option, Option<(u32, u32)>) { + let output = binary().and_then(|binary| { + Command::new(binary) + .arg("-nostdin") + .arg("-i") + .arg(path) + .stdin(Stdio::null()) + .stdout(Stdio::null()) + .stderr(Stdio::piped()) + .output() + .ok() + }); + let Some(output) = output else { + return (None, None); + }; let banner = String::from_utf8_lossy(&output.stderr); + (running_time(&banner), frame_size(&banner)) +} + +/// `Duration: 00:04:03.20, start: ...`. Unparseable for a stream with no known length, which +/// reads as "no timeline" rather than as an error. +fn running_time(banner: &str) -> Option { let clock = banner.split_once("Duration: ")?.1.split(',').next()?; let mut fields = clock.split(':'); - let hours: u32 = fields.next()?.trim().parse().ok()?; - let minutes: u32 = fields.next()?.parse().ok()?; - let seconds: f32 = fields.next()?.parse().ok()?; - Some(hours * 3600 + minutes * 60 + seconds as u32) + let hours: f64 = fields.next()?.trim().parse().ok()?; + let minutes: f64 = fields.next()?.parse().ok()?; + let seconds: f64 = fields.next()?.parse().ok()?; + Some(hours * 3600.0 + minutes * 60.0 + seconds) +} + +/// `Stream #0:0: Video: h264 (avc1 / 0x31637661), yuv420p, 1920x1080 [SAR 1:1 DAR 16:9], ...`: +/// the size is the one comma-separated field that opens with `WxH`. +/// +/// ponytail: the stored size, so a phone clip filmed upright reports itself sideways. Read the +/// stream's rotation if a collection of those turns up. +fn frame_size(banner: &str) -> Option<(u32, u32)> { + let stream = banner.lines().find(|line| line.contains("Video: "))?; + stream.split(", ").find_map(|field| { + let (width, height) = field.split_whitespace().next()?.split_once('x')?; + Some((width.parse().ok()?, height.parse().ok()?)) + }) } fn run(binary: &Path, path: &Path, max_edge: u32, seek: Option<&str>) -> Option { @@ -273,6 +294,19 @@ mod tests { /// What the viewer's scrubber is built on: how long the clip is, and a frame from somewhere /// other than its opening. A `frame_at` that ignored its seek would look right until someone /// noticed every position showing the same picture. + #[test] + fn the_banner_gives_up_a_length_and_a_frame_size() { + let banner = "Input #0, mov,mp4,m4a,3gp,3g2,mj2, from 'clip.mp4':\n \ + Duration: 01:02:03.50, start: 0.000000, bitrate: 2500 kb/s\n \ + Stream #0:0[0x1](und): Video: h264 (High) (avc1 / 0x31637661), \ + yuv420p(tv, bt709, progressive), 1920x1080 [SAR 1:1 DAR 16:9], 30 fps\n"; + assert_eq!(super::running_time(banner), Some(3723.5)); + assert_eq!(super::frame_size(banner), Some((1920, 1080))); + let live = " Duration: N/A, start: 0.000000\n Stream #0:0: Audio: aac, 44100 Hz\n"; + assert_eq!(super::running_time(live), None); + assert_eq!(super::frame_size(live), None); + } + #[test] fn a_clip_reports_its_length_and_yields_a_frame_from_the_middle() { let Some(binary) = super::binary() else { @@ -291,7 +325,11 @@ mod tests { return; } - assert_eq!(media::duration(&path), Some(4), "a four-second clip"); + assert_eq!( + media::probe(&path), + (Some(4.0), Some((320, 240))), + "a four-second clip" + ); let opening = media::frame_at(&path, 128, 0).expect("a frame at the start"); let later = media::frame_at(&path, 128, 3).expect("a frame three seconds in"); assert_ne!( diff --git a/crates/preview/src/pdf.rs b/crates/preview/src/pdf.rs index a7889fa2..38dcc318 100644 --- a/crates/preview/src/pdf.rs +++ b/crates/preview/src/pdf.rs @@ -119,6 +119,18 @@ pub fn page_count(path: &Path) -> Option { usize::try_from(document.pages().len()).ok() } +/// The first page's size in points, rounded: the shape a document has before any page is drawn. +/// `None` when PDFium is absent or the file will not open. +pub fn page_size(path: &Path) -> Option<(u32, u32)> { + let pdfium = locked()?; + let document = pdfium.load_pdf_from_file(path, None).ok()?; + let page = document.pages().get(0).ok()?; + Some(( + page.width().value.round() as u32, + page.height().value.round() as u32, + )) +} + /// The text layer of every page, one page per line break. Empty for a scan that was never OCR'd. /// /// ponytail: holds the PDFium lock for the whole document, so a 500-page PDF stalls gallery @@ -319,6 +331,11 @@ trailer<>"; let path = std::env::temp_dir().join("qrate-pdf-probe.pdf"); std::fs::write(&path, pdf).unwrap(); + assert_eq!( + pdf::page_size(&path), + super::pdfium().map(|_| (200, 100)), + "the MediaBox, or nothing without the library" + ); let rendered = pdf::decode(&path, 128, 0); match (super::pdfium().is_some(), rendered) { (true, Some(page)) => { diff --git a/crates/qrate-export/src/export.rs b/crates/qrate-export/src/export.rs index a39afb28..e21cc328 100644 --- a/crates/qrate-export/src/export.rs +++ b/crates/qrate-export/src/export.rs @@ -4,7 +4,7 @@ //! Everything here takes the grid as plain `headers` + `rows` — the same pair //! `table::save_now` persists — so nothing in this file needs a window or a project handle. -use std::collections::{BTreeMap, HashSet}; +use std::collections::{BTreeMap, HashMap, HashSet}; use std::fs::File; use std::io::Write; use std::path::{Path, PathBuf}; @@ -371,32 +371,11 @@ pub fn zip_to( .map_err(std::io::Error::from)?, )?; - let mut taken: HashSet = HashSet::new(); - for image in images { - let Some(fallback) = image.path.file_name().and_then(|n| n.to_str()) else { + let mut written: HashSet = HashSet::new(); + for (image, name) in images.iter().zip(archive_names(images)) { + let Some(name) = name.filter(|name| written.insert(name.clone())) else { continue; }; - // A file linked from outside the files folder goes under `files/outside/`. - let relative = match image.source_path.as_deref() { - Some(source) if Path::new(source).is_absolute() => format!("outside/{fallback}"), - source => source - .and_then(safe_archive_path) - .unwrap_or_else(|| fallback.to_string()), - }; - // Two folders can hold the same filename, and a zip entry that repeats one silently wins. - let mut name = relative; - for n in 2.. { - if taken.insert(name.clone()) { - break; - } - let stem = Path::new(&name).file_stem().unwrap_or_default(); - let parent = Path::new(&name).parent().unwrap_or_else(|| Path::new("")); - let mut renamed = format!("{}_{n}", stem.to_string_lossy()); - if let Some(ext) = image.path.extension().and_then(|e| e.to_str()) { - renamed = format!("{renamed}.{ext}"); - } - name = parent.join(renamed).to_string_lossy().replace('\\', "/"); - } match File::open(&image.path) { Ok(mut file) => { zip.start_file(format!("files/{name}"), binary)?; @@ -412,6 +391,45 @@ pub fn zip_to( Ok(()) } +/// Where each file lands below `files/`, in the ZIP archive and in a IIIF manifest's addresses +/// alike. `None` for a path with no file name. Rows that link the same file share one name. +pub fn archive_names(images: &[ArchiveFile]) -> Vec> { + let mut taken: HashSet = HashSet::new(); + let mut named: HashMap<&Path, String> = HashMap::new(); + images + .iter() + .map(|image| { + if let Some(name) = named.get(image.path.as_path()) { + return Some(name.clone()); + } + let fallback = image.path.file_name().and_then(|n| n.to_str())?; + // A file linked from outside the files folder goes under `files/outside/`. + let relative = match image.source_path.as_deref() { + Some(source) if Path::new(source).is_absolute() => format!("outside/{fallback}"), + source => source + .and_then(safe_archive_path) + .unwrap_or_else(|| fallback.to_string()), + }; + // Two folders can hold the same filename, and a repeated zip entry silently wins. + let mut name = relative; + for n in 2.. { + if taken.insert(name.clone()) { + break; + } + let stem = Path::new(&name).file_stem().unwrap_or_default(); + let parent = Path::new(&name).parent().unwrap_or_else(|| Path::new("")); + let mut renamed = format!("{}_{n}", stem.to_string_lossy()); + if let Some(ext) = image.path.extension().and_then(|e| e.to_str()) { + renamed = format!("{renamed}.{ext}"); + } + name = parent.join(renamed).to_string_lossy().replace('\\', "/"); + } + named.insert(&image.path, name.clone()); + Some(name) + }) + .collect() +} + fn safe_archive_path(source: &str) -> Option { let path = Path::new(source); if path.is_absolute() || source.trim().is_empty() { @@ -726,6 +744,28 @@ mod tests { ); } + #[test] + fn rows_linking_one_file_share_its_name_and_namesakes_are_told_apart() { + let file = |path: &str, source: &str| ArchiveFile { + path: path.into(), + source_path: Some(source.into()), + }; + let absolute = |name: &str| { + let path = std::env::temp_dir().join(name).join("1.jpg"); + file(&path.to_string_lossy(), &path.to_string_lossy()) + }; + assert_eq!( + super::archive_names(&[ + file("/files/a/1.jpg", "a/1.jpg"), + file("/files/a/1.jpg", "a/1.jpg"), + absolute("x"), + absolute("y"), + ]), + ["a/1.jpg", "a/1.jpg", "outside/1.jpg", "outside/1_2.jpg"] + .map(|name| Some(name.into())) + ); + } + #[test] fn a_file_from_outside_the_files_folder_lands_under_outside() { let dir = std::env::temp_dir().join("qrate-export-zip-outside-test"); diff --git a/crates/qrate-export/src/iiif.rs b/crates/qrate-export/src/iiif.rs new file mode 100644 index 00000000..0f2c2e91 --- /dev/null +++ b/crates/qrate-export/src/iiif.rs @@ -0,0 +1,792 @@ +//! IIIF Presentation API 3: the project as one Manifest, with a Canvas for every row that links a +//! file a viewer can present. +//! +//! Pure like the other writers. The caller measures the files, because the desktop app and a +//! browser reach them differently, and qrate serves none of them: every address is built on the +//! web address the archivist says the export will be published at. + +use std::collections::{HashMap, HashSet}; +use std::path::Path; + +use serde_json::{Map, Value, json}; +use thiserror::Error; + +use crate::ColumnType; +use crate::export::{ArchiveFile, ExportComponent, archive_names}; + +const CONTEXT: &str = "http://iiif.io/api/presentation/3/context.json"; + +/// One row's linked file as it will be published. +#[derive(Clone, Debug, PartialEq)] +pub struct IiifMedia { + /// Its path below `files/`, as [`crate::archive_names`] lays the ZIP archive out. + pub path: String, + /// Pixels: a picture's or a video's own, a document's first page. + pub size: Option<(u32, u32)>, + /// Seconds, for a recording or a video. + pub duration: Option, +} + +/// Each row's linked file under the name it is published by. `extent` answers with a file's pixel +/// size and its running time, whichever it has. +pub fn iiif_media( + linked: &[Option], + extent: impl Fn(&Path) -> (Option<(u32, u32)>, Option), +) -> Vec> { + let files: Vec = linked.iter().flatten().cloned().collect(); + let mut names = archive_names(&files).into_iter(); + linked + .iter() + .map(|file| { + let file = file.as_ref()?; + let path = names.next().flatten()?; + let (size, duration) = extent(&file.path); + Some(IiifMedia { + path, + size, + duration, + }) + }) + .collect() +} + +/// Everything a manifest is built from. `columns` pairs each declared column with its type, and +/// `media` holds one entry per row. +pub struct IiifInput<'a> { + pub base_url: &'a str, + pub label: &'a str, + pub headers: &'a [String], + pub row_ids: &'a [i64], + pub rows: &'a [Vec], + pub structure: &'a [ExportComponent], + pub columns: &'a [(String, ColumnType)], + pub media: &'a [Option], +} + +#[derive(Debug, Error, PartialEq, Eq)] +pub enum IiifError { + #[error("The web address must start with http:// or https:// and have no ? or # in it")] + BaseUrl, + #[error("No row links a file that a IIIF manifest can present")] + Empty, +} + +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum IiifProblem { + /// The row links no file and groups no other rows, so it has nothing to show. + NoFile, + /// The file is of a type IIIF has no way to present. + Unsupported, + /// The file's pixel size or running time could not be read, and a Canvas needs one. + Unmeasured, + /// The rights value is not an address IIIF accepts, so it stays in the metadata only. + Rights, +} + +/// Something the export could not carry over, by the row's position in the grid. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct IiifIssue { + pub row: usize, + pub problem: IiifProblem, +} + +/// The address a manifest is built on, without its trailing slash. +pub fn base_url(value: &str) -> Result<&str, IiifError> { + let base = value.trim().trim_end_matches('/'); + base.strip_prefix("https://") + .or_else(|| base.strip_prefix("http://")) + .filter(|host| { + !host.is_empty() && !host.contains(|c: char| c.is_whitespace() || c == '?' || c == '#') + }) + .map(|_| base) + .ok_or(IiifError::BaseUrl) +} + +/// The manifest, and what it had to leave out. Addresses are stable: a row keeps its Canvas +/// address for as long as it keeps its row id. +pub fn iiif_manifest(input: &IiifInput) -> Result<(Value, Vec), IiifError> { + let base = base_url(input.base_url)?; + + let column = |kind: ColumnType| { + input + .columns + .iter() + .find(|(_, declared)| *declared == kind) + .and_then(|(name, _)| input.headers.iter().position(|header| header == name)) + }; + let (title, identifier, date) = ( + column(ColumnType::Title), + column(ColumnType::Identifier), + column(ColumnType::Date), + ); + let rights = input.headers.iter().position(|header| { + matches!( + header.trim().to_ascii_lowercase().as_str(), + "rights" | "license" | "licence" + ) + }); + let mut children: HashMap> = HashMap::new(); + for component in input.structure { + if let Some(parent) = component.parent_id { + children.entry(parent).or_default().push(component.row_id); + } + } + + let mut issues = Vec::new(); + let mut canvases = Vec::new(); + let mut painted = HashSet::new(); + let mut groups = HashMap::new(); + for (row, (cells, row_id)) in input.rows.iter().zip(input.row_ids).enumerate() { + let cell = |at: Option| { + at.and_then(|i| cells.get(i)) + .map(|value| value.trim()) + .filter(|value| !value.is_empty()) + }; + let media = input.media.get(row).and_then(Option::as_ref); + let label = cell(title) + .or(cell(identifier)) + .or(media.and_then(|media| media.path.rsplit('/').next())) + .map_or_else(|| format!("Row {}", row + 1), str::to_owned); + let metadata: Vec = input + .headers + .iter() + .zip(cells) + .filter(|(_, value)| !value.trim().is_empty()) + .map(|(header, value)| json!({ "label": text(header), "value": text(value) })) + .collect(); + let grouping = children.contains_key(row_id); + if grouping { + groups.insert(*row_id, (label.clone(), metadata.clone())); + } + let mut issue = |problem| issues.push(IiifIssue { row, problem }); + + let Some(media) = media else { + if !grouping { + issue(IiifProblem::NoFile); + } + continue; + }; + let Some((kind, format)) = content_type(&media.path) else { + issue(IiifProblem::Unsupported); + continue; + }; + let size = media + .size + .filter(|(width, height)| *width > 0 && *height > 0); + let duration = media + .duration + .filter(|seconds| seconds.is_finite() && *seconds > 0.0); + let (size, duration) = match kind { + "Image" | "Text" => (size, None), + "Sound" => (None, duration), + _ => (size, duration), + }; + let measured = match kind { + "Image" | "Text" => size.is_some(), + _ => duration.is_some(), + }; + if !measured { + issue(IiifProblem::Unmeasured); + continue; + } + + let id = format!("{base}/canvas/{row_id}"); + let mut body = Map::new(); + body.insert( + "id".into(), + format!("{base}/files/{}", encode_path(&media.path)).into(), + ); + body.insert("type".into(), kind.into()); + body.insert("format".into(), format.into()); + let mut canvas = Map::new(); + canvas.insert("id".into(), id.clone().into()); + canvas.insert("type".into(), "Canvas".into()); + canvas.insert("label".into(), text(&label)); + for resource in [&mut canvas, &mut body] { + if let Some((width, height)) = size { + resource.insert("width".into(), width.into()); + resource.insert("height".into(), height.into()); + } + if let Some(seconds) = duration { + resource.insert("duration".into(), seconds.into()); + } + } + if let Some(instant) = cell(date).and_then(nav_date) { + canvas.insert("navDate".into(), instant.into()); + } + match cell(rights).map(rights_uri) { + Some(Some(uri)) => { + canvas.insert("rights".into(), uri.into()); + } + Some(None) => issue(IiifProblem::Rights), + None => {} + } + if !metadata.is_empty() { + canvas.insert("metadata".into(), metadata.into()); + } + canvas.insert( + "items".into(), + json!([{ + "id": format!("{id}/page"), + "type": "AnnotationPage", + "items": [{ + "id": format!("{id}/painting"), + "type": "Annotation", + "motivation": "painting", + "body": body, + "target": id, + }], + }]), + ); + canvases.push(Value::Object(canvas)); + painted.insert(*row_id); + } + if canvases.is_empty() { + return Err(IiifError::Empty); + } + + let ranges = Ranges { + base, + children: &children, + painted: &painted, + groups: &groups, + }; + let mut seen = HashSet::new(); + let structures: Vec = input + .structure + .iter() + .filter(|component| component.parent_id.is_none()) + .filter_map(|component| ranges.build(component.row_id, &mut seen)) + .collect(); + + let mut manifest = Map::new(); + manifest.insert("@context".into(), CONTEXT.into()); + manifest.insert("id".into(), format!("{base}/manifest.json").into()); + manifest.insert("type".into(), "Manifest".into()); + manifest.insert( + "label".into(), + text(match input.label.trim() { + "" => "Untitled", + label => label, + }), + ); + manifest.insert("items".into(), canvases.into()); + if !structures.is_empty() { + manifest.insert("structures".into(), structures.into()); + } + Ok((Value::Object(manifest), issues)) +} + +/// The archival arrangement as nested Ranges: the viewer's table of contents. +struct Ranges<'a> { + base: &'a str, + children: &'a HashMap>, + painted: &'a HashSet, + groups: &'a HashMap)>, +} + +impl Ranges<'_> { + /// `None` for a row that groups nothing a viewer can show, since a Range may not be empty. + /// `seen` stops a project whose structure loops back on itself from recursing forever. + fn build(&self, row_id: i64, seen: &mut HashSet) -> Option { + let below = self.children.get(&row_id)?; + if !seen.insert(row_id) { + return None; + } + let canvas = |row_id: i64| { + self.painted.contains(&row_id).then( + || json!({ "id": format!("{}/canvas/{row_id}", self.base), "type": "Canvas" }), + ) + }; + let mut items: Vec = canvas(row_id).into_iter().collect(); + for child in below { + items.extend(match self.children.contains_key(child) { + true => self.build(*child, seen), + false => canvas(*child), + }); + } + if items.is_empty() { + return None; + } + let mut range = Map::new(); + range.insert("id".into(), format!("{}/range/{row_id}", self.base).into()); + range.insert("type".into(), "Range".into()); + match self.groups.get(&row_id) { + Some((label, metadata)) => { + range.insert("label".into(), text(label)); + if !metadata.is_empty() { + range.insert("metadata".into(), metadata.clone().into()); + } + } + None => { + range.insert("label".into(), text(&row_id.to_string())); + } + } + range.insert("items".into(), items.into()); + Some(Value::Object(range)) + } +} + +/// What an export could not carry over, in words, or nothing when it carried everything. +pub fn iiif_summary(issues: &[IiifIssue]) -> String { + let count = |problem| { + issues + .iter() + .filter(|issue| issue.problem == problem) + .count() + }; + let rows = |n: usize| format!("{n} row{}", if n == 1 { "" } else { "s" }); + let mut parts = Vec::new(); + for (problem, why) in [ + (IiifProblem::NoFile, "with no linked file"), + ( + IiifProblem::Unsupported, + "with a file type IIIF cannot present", + ), + ( + IiifProblem::Unmeasured, + "whose file's size or length could not be read", + ), + ] { + match count(problem) { + 0 => {} + n => parts.push(format!("left out {} {why}", rows(n))), + } + } + match count(IiifProblem::Rights) { + 0 => {} + n => parts.push(format!( + "kept {} rights value{} as metadata only, because IIIF accepts just Creative Commons \ + and RightsStatements.org addresses there", + n, + if n == 1 { "" } else { "s" } + )), + } + parts.join("; ") +} + +/// A value in no particular language, which is all a spreadsheet cell says about itself. +fn text(value: &str) -> Value { + json!({ "none": [value] }) +} + +/// The IIIF content type and the media type of a file, by its extension: every kind +/// `preview::extent` can measure. A PDF is `Text`, which is how viewers that show documents expect +/// to be handed one. +fn content_type(path: &str) -> Option<(&'static str, &'static str)> { + let extension = path.rsplit_once('.')?.1.to_ascii_lowercase(); + Some(match extension.as_str() { + "jpg" | "jpeg" => ("Image", "image/jpeg"), + "png" => ("Image", "image/png"), + "gif" => ("Image", "image/gif"), + "webp" => ("Image", "image/webp"), + "bmp" => ("Image", "image/bmp"), + "tif" | "tiff" => ("Image", "image/tiff"), + "avif" => ("Image", "image/avif"), + "jp2" => ("Image", "image/jp2"), + "jpf" | "jpx" => ("Image", "image/jpx"), + "j2k" => ("Image", "image/j2c"), + "heic" => ("Image", "image/heic"), + "heif" => ("Image", "image/heif"), + "mp4" | "m4v" => ("Video", "video/mp4"), + "mov" => ("Video", "video/quicktime"), + "webm" => ("Video", "video/webm"), + "mkv" => ("Video", "video/x-matroska"), + "avi" => ("Video", "video/x-msvideo"), + "mpg" | "mpeg" => ("Video", "video/mpeg"), + "wmv" => ("Video", "video/x-ms-wmv"), + "flv" => ("Video", "video/x-flv"), + "m2ts" => ("Video", "video/mp2t"), + "mp3" => ("Sound", "audio/mpeg"), + "wav" => ("Sound", "audio/wav"), + "flac" => ("Sound", "audio/flac"), + "ogg" | "oga" => ("Sound", "audio/ogg"), + "m4a" => ("Sound", "audio/mp4"), + "aac" => ("Sound", "audio/aac"), + "aif" | "aiff" => ("Sound", "audio/aiff"), + "pdf" => ("Text", "application/pdf"), + _ => return None, + }) +} + +/// A path as it appears in a web address, each segment escaped. +fn encode_path(path: &str) -> String { + path.bytes() + .map(|byte| match byte { + b'A'..=b'Z' | b'a'..=b'z' | b'0'..=b'9' | b'-' | b'.' | b'_' | b'~' | b'/' => { + char::from(byte).to_string() + } + _ => format!("%{byte:02X}"), + }) + .collect() +} + +/// A plain year, month or day as the instant it begins, which is what a viewer's timeline wants. +/// A range or an approximate date names no single instant, so it stays in the metadata only. +fn nav_date(value: &str) -> Option { + let number = |part: &str, digits: usize| { + (part.len() == digits && part.bytes().all(|byte| byte.is_ascii_digit())) + .then(|| part.parse::().ok()) + .flatten() + }; + let mut parts = value.split('-'); + let year = number(parts.next()?, 4).filter(|year| *year > 0)?; + let month = parts.next().map_or(Some(1), |part| number(part, 2))?; + let day = parts.next().map_or(Some(1), |part| number(part, 2))?; + let leap = year % 4 == 0 && (year % 100 != 0 || year % 400 == 0); + let days = match month { + 2 if leap => 29, + 2 => 28, + 4 | 6 | 9 | 11 => 30, + 1..=12 => 31, + _ => return None, + }; + (parts.next().is_none() && (1..=days).contains(&day)) + .then(|| format!("{year:04}-{month:02}-{day:02}T00:00:00Z")) +} + +/// A Creative Commons or RightsStatements.org address in the `http://` form IIIF asks for, which +/// is how both publish their identifiers even though their pages are served over `https://`. +fn rights_uri(value: &str) -> Option { + let rest = value + .strip_prefix("https://") + .or_else(|| value.strip_prefix("http://"))?; + [ + "creativecommons.org/licenses/", + "creativecommons.org/publicdomain/", + "rightsstatements.org/vocab/", + ] + .iter() + .any(|known| rest.len() > known.len() && rest.starts_with(known)) + .then(|| format!("http://{rest}")) + .filter(|uri| !uri.contains(char::is_whitespace)) +} + +#[cfg(test)] +mod tests { + use super::{ + IiifError, IiifInput, IiifIssue, IiifMedia, IiifProblem, content_type, iiif_manifest, + iiif_media, iiif_summary, nav_date, rights_uri, + }; + use crate::ColumnType; + use crate::export::{ArchiveFile, ExportComponent}; + + #[test] + fn each_row_keeps_its_own_file_when_some_rows_link_none() { + let file = |path: &str| { + Some(ArchiveFile { + path: path.into(), + source_path: Some(path.into()), + }) + }; + let media = iiif_media(&[None, file("a/1.jpg"), None, file("b/1.jpg")], |path| { + (Some((path.to_string_lossy().len() as u32, 1)), None) + }); + assert_eq!( + media + .iter() + .map(|media| media.as_ref().map(|media| media.path.as_str())) + .collect::>(), + [None, Some("a/1.jpg"), None, Some("b/1.jpg")] + ); + assert_eq!(media[3].as_ref().unwrap().size, Some((7, 1))); + } + + fn image(path: &str) -> Option { + Some(IiifMedia { + path: path.into(), + size: Some((4000, 3000)), + duration: None, + }) + } + + struct Fixture { + headers: Vec, + rows: Vec>, + structure: Vec, + columns: Vec<(String, ColumnType)>, + media: Vec>, + } + + impl Fixture { + fn input<'a>(&'a self, base_url: &'a str) -> IiifInput<'a> { + IiifInput { + base_url, + label: "Harbour photographs", + headers: &self.headers, + row_ids: &[10, 11, 12, 13, 14, 15, 16], + rows: &self.rows, + structure: &self.structure, + columns: &self.columns, + media: &self.media, + } + } + } + + /// A series holding two photographs, a recording and a video, then two loose rows: one that + /// links nothing and one that links a document. + fn fixture() -> Fixture { + let component = |row_id, parent_id: Option, level: &str| ExportComponent { + row_id, + parent_id, + level_key: level.into(), + source_path: None, + }; + let row = |cells: [&str; 5]| cells.map(String::from).to_vec(); + Fixture { + headers: ["Digital ID", "Title", "Taken", "Rights", "File"] + .map(String::from) + .to_vec(), + rows: vec![ + row(["S1", "Harbour series", "1943/1950", "", ""]), + row([ + "1", + "First photo", + "1943-06", + "https://creativecommons.org/licenses/by/4.0/", + "a/first photo.jpg", + ]), + row(["2", "", "circa 1944", "ask the donor", "a/2.tif"]), + row(["3", "Interview", "1950", "", "tape.mp3"]), + row(["4", "Launch day", "", "", "launch.mp4"]), + row(["5", "Loose note", "", "", ""]), + row(["6", "Finding aid", "", "", "finding aid.pdf"]), + ], + structure: vec![ + component(10, None, "series"), + component(11, Some(10), "item"), + component(12, Some(10), "item"), + component(13, Some(10), "item"), + component(14, Some(10), "item"), + ], + columns: vec![ + ("Digital ID".into(), ColumnType::Identifier), + ("Title".into(), ColumnType::Title), + ("Taken".into(), ColumnType::Date), + ("File".into(), ColumnType::Filename), + ], + media: vec![ + None, + image("a/first photo.jpg"), + image("a/2.tif"), + Some(IiifMedia { + path: "tape.mp3".into(), + size: None, + duration: Some(1845.5), + }), + Some(IiifMedia { + path: "launch.mp4".into(), + size: Some((1920, 1080)), + duration: Some(62.0), + }), + None, + Some(IiifMedia { + path: "finding aid.pdf".into(), + size: Some((612, 792)), + duration: None, + }), + ], + } + } + + /// The golden file is the one checked against the IIIF Presentation validator; when this + /// output changes on purpose, validate the new file before committing it. + #[test] + fn a_project_becomes_one_manifest_with_a_canvas_per_linked_file() { + let fixture = fixture(); + let (manifest, issues) = + iiif_manifest(&fixture.input("https://example.org/iiif/harbour/")).unwrap(); + let golden: serde_json::Value = + serde_json::from_str(include_str!("../tests/golden/catalog.iiif.json")).unwrap(); + assert_eq!(manifest, golden); + + assert_eq!( + manifest["id"], + "https://example.org/iiif/harbour/manifest.json" + ); + let canvases = manifest["items"].as_array().unwrap(); + assert_eq!(canvases.len(), 5); + let body = &canvases[0]["items"][0]["items"][0]["body"]; + assert_eq!( + body["id"], + "https://example.org/iiif/harbour/files/a/first%20photo.jpg" + ); + assert_eq!(canvases[0]["navDate"], "1943-06-01T00:00:00Z"); + assert_eq!( + canvases[0]["rights"], + "http://creativecommons.org/licenses/by/4.0/" + ); + // An untitled row is labelled by its identifier, and a date that names no instant and a + // rights note that is no licence address stay where the archivist wrote them. + assert_eq!(canvases[1]["label"]["none"][0], "2"); + assert!(canvases[1].get("navDate").is_none()); + assert!(canvases[1].get("rights").is_none()); + assert_eq!(canvases[2]["duration"], 1845.5); + assert!(canvases[2].get("width").is_none()); + assert_eq!(canvases[3]["width"], 1920); + assert_eq!(canvases[4]["items"][0]["items"][0]["body"]["type"], "Text"); + + let series = &manifest["structures"][0]; + assert_eq!(series["label"]["none"][0], "Harbour series"); + assert_eq!(series["items"].as_array().unwrap().len(), 4); + assert_eq!( + issues, + [ + IiifIssue { + row: 2, + problem: IiifProblem::Rights + }, + IiifIssue { + row: 5, + problem: IiifProblem::NoFile + }, + ] + ); + } + + #[test] + fn a_file_that_cannot_be_presented_is_reported_instead_of_written_malformed() { + let mut fixture = fixture(); + fixture.media[1] = Some(IiifMedia { + path: "a/first photo.jpg".into(), + size: None, + duration: None, + }); + fixture.media[2] = image("a/notes.docx"); + fixture.media[4] = Some(IiifMedia { + path: "launch.mp4".into(), + size: Some((1920, 1080)), + duration: None, + }); + let (manifest, issues) = iiif_manifest(&fixture.input("https://example.org")).unwrap(); + assert_eq!(manifest["items"].as_array().unwrap().len(), 2); + assert_eq!( + issues.iter().map(|issue| issue.problem).collect::>(), + [ + IiifProblem::Unmeasured, + IiifProblem::Unsupported, + IiifProblem::Unmeasured, + IiifProblem::NoFile, + ] + ); + assert_eq!( + iiif_summary(&issues), + "left out 1 row with no linked file; left out 1 row with a file type IIIF cannot \ + present; left out 2 rows whose file's size or length could not be read" + ); + assert_eq!(iiif_summary(&[]), ""); + } + + #[test] + fn an_address_that_is_not_on_the_web_or_a_project_with_nothing_to_show_is_refused() { + let mut fixture = fixture(); + for base in [ + "", + "example.org", + "ftp://example.org", + "https://", + "https://x.org/a#b", + ] { + assert_eq!( + iiif_manifest(&fixture.input(base)).unwrap_err(), + IiifError::BaseUrl, + "{base}" + ); + } + fixture.media = vec![None; 7]; + assert_eq!( + iiif_manifest(&fixture.input("https://example.org")).unwrap_err(), + IiifError::Empty + ); + } + + #[test] + fn a_structure_that_loops_back_on_itself_still_ends() { + let mut fixture = fixture(); + fixture.structure.push(ExportComponent { + row_id: 10, + parent_id: Some(11), + level_key: "series".into(), + source_path: None, + }); + let (manifest, _) = iiif_manifest(&fixture.input("https://example.org")).unwrap(); + let series = manifest["structures"][0]["items"].as_array().unwrap(); + assert_eq!( + series[0]["type"], "Range", + "the photo that also groups rows" + ); + assert_eq!(series[0]["items"].as_array().unwrap().len(), 1); + } + + #[test] + fn a_group_with_nothing_to_show_yields_no_range() { + let mut fixture = fixture(); + fixture.structure.truncate(1); + fixture.structure.push(ExportComponent { + row_id: 15, + parent_id: Some(10), + level_key: "item".into(), + source_path: None, + }); + let (manifest, _) = iiif_manifest(&fixture.input("https://example.org")).unwrap(); + assert!(manifest.get("structures").is_none()); + } + + #[test] + fn only_a_whole_real_date_becomes_a_navigation_date() { + assert_eq!(nav_date("1943").as_deref(), Some("1943-01-01T00:00:00Z")); + assert_eq!( + nav_date("2024-02-29").as_deref(), + Some("2024-02-29T00:00:00Z") + ); + for not_an_instant in [ + "1943/1950", + "1943~", + "194X", + "2023-02-29", + "1943-13", + "43", + "1943-6", + ] { + assert_eq!(nav_date(not_an_instant), None, "{not_an_instant}"); + } + } + + #[test] + fn every_kind_of_file_preview_measures_has_a_type() { + for still in ["a.JP2", "a.jpf", "a.jpx", "a.j2k", "a.heic", "a.avif"] { + assert_eq!( + content_type(still).map(|(kind, _)| kind), + Some("Image"), + "{still}" + ); + } + for video in ["a.wmv", "a.flv", "a.m2ts", "a.mov"] { + assert_eq!( + content_type(video).map(|(kind, _)| kind), + Some("Video"), + "{video}" + ); + } + assert_eq!(content_type("a.docx"), None); + } + + #[test] + fn rights_are_kept_only_as_the_addresses_iiif_accepts() { + assert_eq!( + rights_uri("https://rightsstatements.org/vocab/InC/1.0/").as_deref(), + Some("http://rightsstatements.org/vocab/InC/1.0/") + ); + assert_eq!( + rights_uri("http://creativecommons.org/publicdomain/zero/1.0/").as_deref(), + Some("http://creativecommons.org/publicdomain/zero/1.0/") + ); + for refused in [ + "CC BY 4.0", + "https://example.org/licence", + "https://creativecommons.org/licenses/", + ] { + assert_eq!(rights_uri(refused), None, "{refused}"); + } + } +} diff --git a/crates/qrate-export/src/lib.rs b/crates/qrate-export/src/lib.rs index 13d60530..5f19bb68 100644 --- a/crates/qrate-export/src/lib.rs +++ b/crates/qrate-export/src/lib.rs @@ -6,6 +6,7 @@ pub mod columns; pub mod description; pub mod export; pub mod filenames; +pub mod iiif; pub mod notes; pub mod photos; #[cfg(feature = "wasm")] @@ -13,8 +14,13 @@ pub mod wasm; pub use columns::ColumnType; pub use export::{ - ArchiveFile, CSL_FIELDS, CslMapping, CsvOptions, ExportComponent, csl_items, csv_bytes, - derive_csl_mapping, jsonld_hierarchy_value, project_structure_columns, xlsx_bytes, zip_to, + ArchiveFile, CSL_FIELDS, CslMapping, CsvOptions, ExportComponent, archive_names, csl_items, + csv_bytes, derive_csl_mapping, jsonld_hierarchy_value, project_structure_columns, xlsx_bytes, + zip_to, +}; +pub use iiif::{ + IiifError, IiifInput, IiifIssue, IiifMedia, IiifProblem, iiif_manifest, iiif_media, + iiif_summary, }; pub use notes::{ProjectNote, SheetNote, read_project_notes, sheet_note_request_body, sheet_notes}; pub use photos::{PhotoIndex, relative_to}; diff --git a/crates/qrate-export/src/wasm.rs b/crates/qrate-export/src/wasm.rs index 42688bab..bfa19db7 100644 --- a/crates/qrate-export/src/wasm.rs +++ b/crates/qrate-export/src/wasm.rs @@ -1,12 +1,17 @@ //! Browser adapter for the shared project reader and format writers. +use std::collections::HashMap; use std::io::Cursor; +use std::path::PathBuf; use rusqlite::{Connection, MAIN_DB, OptionalExtension}; use wasm_bindgen::prelude::*; -use crate::export::{CslMapping, ExportComponent}; -use crate::{ColumnType, QRATE_APPLICATION_ID, QRATE_SCHEMA_VERSION, read_dataset}; +use crate::export::{ArchiveFile, CslMapping, ExportComponent}; +use crate::{ + ColumnType, IiifError, IiifInput, PhotoIndex, QRATE_APPLICATION_ID, QRATE_SCHEMA_VERSION, + read_dataset, +}; #[derive(serde::Deserialize)] struct DescriptionLevel { @@ -14,6 +19,36 @@ struct DescriptionLevel { label: String, } +/// A file the page found in the folder the archivist picked, measured by the browser. A zero +/// stands for a size or a running time the file does not have. +#[derive(serde::Deserialize)] +struct MeasuredFile { + path: String, + #[serde(default)] + width: u32, + #[serde(default)] + height: u32, + #[serde(default)] + duration: f64, +} + +/// A IIIF manifest, and what it could not carry over, in the words the desktop app uses. +#[wasm_bindgen] +pub struct IiifExport { + manifest: Vec, + summary: String, +} + +#[wasm_bindgen] +impl IiifExport { + pub fn manifest(&self) -> Vec { + self.manifest.clone() + } + pub fn summary(&self) -> String { + self.summary.clone() + } +} + fn error(code: &str, detail: impl std::fmt::Display) -> JsError { JsError::new(&format!("{code}: {detail}")) } @@ -27,6 +62,7 @@ pub struct Project { structure: Vec, types: Vec<(String, String)>, mapping: Option, + iiif_base: Option, sheet_notes: Vec, } @@ -128,6 +164,7 @@ impl Project { structure, types, mapping, + iiif_base: setting("iiif_base_url")?, sheet_notes, }) } @@ -181,6 +218,70 @@ impl Project { serde_json::to_vec_pretty(&crate::csl_items(&self.headers, &self.rows, &mapping)) .map_err(|e| error("write", e)) } + /// The web address the desktop app last published this project's manifest at. + pub fn iiif_base(&self) -> Option { + self.iiif_base.clone() + } + /// `files` is every file in the folder the archivist picked, as `{path, width, height, + /// duration}` with paths relative to that folder. + pub fn to_iiif( + &self, + base_url: &str, + label: &str, + files: JsValue, + ) -> Result { + let files: Vec = + serde_wasm_bindgen::from_value(files).map_err(|e| error("files", e))?; + let measured: HashMap = files + .iter() + .map(|file| (PathBuf::from(&file.path), file)) + .collect(); + let index = PhotoIndex::from_paths(measured.keys().cloned()); + let columns: Vec<(String, ColumnType)> = self + .types + .iter() + .map(|(name, ty)| (name.clone(), ColumnType::from_declared(ty))) + .collect(); + let declared: Vec = columns + .iter() + .filter(|(_, kind)| *kind == ColumnType::Filename) + .map(|(name, _)| name.clone()) + .collect(); + let linked: Vec> = self + .rows + .iter() + .map(|row| { + let path = index.resolve_row(&self.headers, &declared, row)?; + Some(ArchiveFile { + source_path: Some(path.to_string_lossy().replace('\\', "/")), + path, + }) + }) + .collect(); + let media = crate::iiif_media(&linked, |path| { + measured.get(path).map_or((None, None), |file| { + (Some((file.width, file.height)), Some(file.duration)) + }) + }); + let (manifest, issues) = crate::iiif_manifest(&IiifInput { + base_url, + label, + headers: &self.headers, + row_ids: &self.row_ids, + rows: &self.rows, + structure: &self.structure, + columns: &columns, + media: &media, + }) + .map_err(|err| match err { + IiifError::BaseUrl => error("address", err), + IiifError::Empty => error("nothing", err), + })?; + Ok(IiifExport { + manifest: serde_json::to_vec_pretty(&manifest).map_err(|e| error("write", e))?, + summary: crate::iiif_summary(&issues), + }) + } pub fn sheet_values(&self) -> JsValue { let mut values = vec![&self.headers]; values.extend(self.rows.iter()); diff --git a/crates/qrate-export/tests/golden/catalog.iiif.json b/crates/qrate-export/tests/golden/catalog.iiif.json new file mode 100644 index 00000000..a5e03e1b --- /dev/null +++ b/crates/qrate-export/tests/golden/catalog.iiif.json @@ -0,0 +1,480 @@ +{ + "@context": "http://iiif.io/api/presentation/3/context.json", + "id": "https://example.org/iiif/harbour/manifest.json", + "type": "Manifest", + "label": { + "none": [ + "Harbour photographs" + ] + }, + "items": [ + { + "id": "https://example.org/iiif/harbour/canvas/11", + "type": "Canvas", + "label": { + "none": [ + "First photo" + ] + }, + "width": 4000, + "height": 3000, + "navDate": "1943-06-01T00:00:00Z", + "rights": "http://creativecommons.org/licenses/by/4.0/", + "metadata": [ + { + "label": { + "none": [ + "Digital ID" + ] + }, + "value": { + "none": [ + "1" + ] + } + }, + { + "label": { + "none": [ + "Title" + ] + }, + "value": { + "none": [ + "First photo" + ] + } + }, + { + "label": { + "none": [ + "Taken" + ] + }, + "value": { + "none": [ + "1943-06" + ] + } + }, + { + "label": { + "none": [ + "Rights" + ] + }, + "value": { + "none": [ + "https://creativecommons.org/licenses/by/4.0/" + ] + } + }, + { + "label": { + "none": [ + "File" + ] + }, + "value": { + "none": [ + "a/first photo.jpg" + ] + } + } + ], + "items": [ + { + "id": "https://example.org/iiif/harbour/canvas/11/page", + "type": "AnnotationPage", + "items": [ + { + "id": "https://example.org/iiif/harbour/canvas/11/painting", + "type": "Annotation", + "motivation": "painting", + "body": { + "id": "https://example.org/iiif/harbour/files/a/first%20photo.jpg", + "type": "Image", + "format": "image/jpeg", + "width": 4000, + "height": 3000 + }, + "target": "https://example.org/iiif/harbour/canvas/11" + } + ] + } + ] + }, + { + "id": "https://example.org/iiif/harbour/canvas/12", + "type": "Canvas", + "label": { + "none": [ + "2" + ] + }, + "width": 4000, + "height": 3000, + "metadata": [ + { + "label": { + "none": [ + "Digital ID" + ] + }, + "value": { + "none": [ + "2" + ] + } + }, + { + "label": { + "none": [ + "Taken" + ] + }, + "value": { + "none": [ + "circa 1944" + ] + } + }, + { + "label": { + "none": [ + "Rights" + ] + }, + "value": { + "none": [ + "ask the donor" + ] + } + }, + { + "label": { + "none": [ + "File" + ] + }, + "value": { + "none": [ + "a/2.tif" + ] + } + } + ], + "items": [ + { + "id": "https://example.org/iiif/harbour/canvas/12/page", + "type": "AnnotationPage", + "items": [ + { + "id": "https://example.org/iiif/harbour/canvas/12/painting", + "type": "Annotation", + "motivation": "painting", + "body": { + "id": "https://example.org/iiif/harbour/files/a/2.tif", + "type": "Image", + "format": "image/tiff", + "width": 4000, + "height": 3000 + }, + "target": "https://example.org/iiif/harbour/canvas/12" + } + ] + } + ] + }, + { + "id": "https://example.org/iiif/harbour/canvas/13", + "type": "Canvas", + "label": { + "none": [ + "Interview" + ] + }, + "duration": 1845.5, + "navDate": "1950-01-01T00:00:00Z", + "metadata": [ + { + "label": { + "none": [ + "Digital ID" + ] + }, + "value": { + "none": [ + "3" + ] + } + }, + { + "label": { + "none": [ + "Title" + ] + }, + "value": { + "none": [ + "Interview" + ] + } + }, + { + "label": { + "none": [ + "Taken" + ] + }, + "value": { + "none": [ + "1950" + ] + } + }, + { + "label": { + "none": [ + "File" + ] + }, + "value": { + "none": [ + "tape.mp3" + ] + } + } + ], + "items": [ + { + "id": "https://example.org/iiif/harbour/canvas/13/page", + "type": "AnnotationPage", + "items": [ + { + "id": "https://example.org/iiif/harbour/canvas/13/painting", + "type": "Annotation", + "motivation": "painting", + "body": { + "id": "https://example.org/iiif/harbour/files/tape.mp3", + "type": "Sound", + "format": "audio/mpeg", + "duration": 1845.5 + }, + "target": "https://example.org/iiif/harbour/canvas/13" + } + ] + } + ] + }, + { + "id": "https://example.org/iiif/harbour/canvas/14", + "type": "Canvas", + "label": { + "none": [ + "Launch day" + ] + }, + "width": 1920, + "height": 1080, + "duration": 62.0, + "metadata": [ + { + "label": { + "none": [ + "Digital ID" + ] + }, + "value": { + "none": [ + "4" + ] + } + }, + { + "label": { + "none": [ + "Title" + ] + }, + "value": { + "none": [ + "Launch day" + ] + } + }, + { + "label": { + "none": [ + "File" + ] + }, + "value": { + "none": [ + "launch.mp4" + ] + } + } + ], + "items": [ + { + "id": "https://example.org/iiif/harbour/canvas/14/page", + "type": "AnnotationPage", + "items": [ + { + "id": "https://example.org/iiif/harbour/canvas/14/painting", + "type": "Annotation", + "motivation": "painting", + "body": { + "id": "https://example.org/iiif/harbour/files/launch.mp4", + "type": "Video", + "format": "video/mp4", + "width": 1920, + "height": 1080, + "duration": 62.0 + }, + "target": "https://example.org/iiif/harbour/canvas/14" + } + ] + } + ] + }, + { + "id": "https://example.org/iiif/harbour/canvas/16", + "type": "Canvas", + "label": { + "none": [ + "Finding aid" + ] + }, + "width": 612, + "height": 792, + "metadata": [ + { + "label": { + "none": [ + "Digital ID" + ] + }, + "value": { + "none": [ + "6" + ] + } + }, + { + "label": { + "none": [ + "Title" + ] + }, + "value": { + "none": [ + "Finding aid" + ] + } + }, + { + "label": { + "none": [ + "File" + ] + }, + "value": { + "none": [ + "finding aid.pdf" + ] + } + } + ], + "items": [ + { + "id": "https://example.org/iiif/harbour/canvas/16/page", + "type": "AnnotationPage", + "items": [ + { + "id": "https://example.org/iiif/harbour/canvas/16/painting", + "type": "Annotation", + "motivation": "painting", + "body": { + "id": "https://example.org/iiif/harbour/files/finding%20aid.pdf", + "type": "Text", + "format": "application/pdf", + "width": 612, + "height": 792 + }, + "target": "https://example.org/iiif/harbour/canvas/16" + } + ] + } + ] + } + ], + "structures": [ + { + "id": "https://example.org/iiif/harbour/range/10", + "type": "Range", + "label": { + "none": [ + "Harbour series" + ] + }, + "metadata": [ + { + "label": { + "none": [ + "Digital ID" + ] + }, + "value": { + "none": [ + "S1" + ] + } + }, + { + "label": { + "none": [ + "Title" + ] + }, + "value": { + "none": [ + "Harbour series" + ] + } + }, + { + "label": { + "none": [ + "Taken" + ] + }, + "value": { + "none": [ + "1943/1950" + ] + } + } + ], + "items": [ + { + "id": "https://example.org/iiif/harbour/canvas/11", + "type": "Canvas" + }, + { + "id": "https://example.org/iiif/harbour/canvas/12", + "type": "Canvas" + }, + { + "id": "https://example.org/iiif/harbour/canvas/13", + "type": "Canvas" + }, + { + "id": "https://example.org/iiif/harbour/canvas/14", + "type": "Canvas" + } + ] + } + ] +} diff --git a/docs/export-and-sync.md b/docs/export-and-sync.md index b9919018..8d2220be 100644 --- a/docs/export-and-sync.md +++ b/docs/export-and-sync.md @@ -10,6 +10,8 @@ Export the project from **File ▸ Export** as one of: - **JSON-LD**. Each row records the row it is part of, so the archival arrangement survives. See [Groups](grid.md#groups). - **Zotero (CSL-JSON)** +- **IIIF Manifest**, which describes the project and its linked files for IIIF viewers such as + Mirador and Universal Viewer. See [IIIF manifests](#iiif-manifests). - **ZIP Archive**, which bundles the data as CSV and JSON-LD with the linked files it points to. Files imported from folders keep the folder paths they were imported from. @@ -43,6 +45,59 @@ Export always reads every row, regardless of any active filter. See No qrate installed? [Convert a project in your browser](https://qrate.dvnl.work/convert). +## IIIF manifests + +**File ▸ Export ▸ IIIF Manifest…** writes the project as one +[IIIF Presentation API 3.0](https://iiif.io/api/presentation/3.0/) manifest. Every row that +links a file becomes one item in it, in table order, and the archival arrangement becomes the +viewer's table of contents. See [Groups](grid.md#groups). + +A manifest names everything by web address, and qrate does not host your files. So the export +first asks for the web address of the folder you will publish the manifest in, such as +`https://example.org/iiif/harbour`. qrate remembers the address in the project. To publish: + +1. Export the manifest. Keep the name `manifest.json`, because the manifest's own address ends + in it. +2. Export a **ZIP Archive** and unpack it. Its `files` folder holds every linked file at the + path the manifest expects. +3. Upload `manifest.json` and the `files` folder into the folder at that web address. + +A viewer on another site can load the files only if your web server allows it, which servers +call CORS. + +What goes into the manifest: + +| In the project | In the manifest | +|---|---| +| The project's name | The manifest's label | +| The **Title** column, or else the **Identifier** column | Each item's label | +| Every cell that is not empty | Each item's metadata, under the column's name | +| A **Date** column holding one year, month, or day, such as `1943` or `1943-06-12` | The item's navigation date | +| A column named **Rights** or **License** holding a Creative Commons or RightsStatements.org web address | The item's rights | +| A row that groups other rows | A section in the table of contents | + +A date range, an approximate date, or a rights note in words stays in the metadata, where a +viewer still shows it. + +Which files a manifest can present: + +- **Pictures**: JPEG, PNG, GIF, WebP, BMP, TIFF, AVIF, HEIC, and JPEG 2000. Most web browsers + cannot show TIFF, HEIC, or JPEG 2000 directly, so convert those for viewing or publish them + through an image server. +- **Recordings and video**, with their running time. Video needs ffmpeg. +- **PDF documents**, at the size of their first page. This needs PDFium. Not every + viewer shows documents. Mirador, for one, shows only their description. + +ffmpeg and PDFium are the same parts that draw previews. See +[Viewing a file](files-and-photos.md#viewing-a-file). + +When the export finishes, its notice says how many rows went in and why any were left out: a +row with no linked file, a file type IIIF cannot present, or a file whose size or length could +not be read. qrate leaves those rows out instead of writing a manifest that viewers reject. + +qrate writes manifests for the Presentation API only. It does not serve images, so it is not a +IIIF Image API server, and the manifest does not refer to one. + ## Google Sheets Google Sheets export and sync is off by default. Turn it on in **Settings ▸ Google**. Until