How to build dataset views with View and Operation¶
Problem statement¶
Analysis rarely runs on a dataset exactly as it ships. You need a curated version of it: only certain classes, a balanced class distribution, a capped size, a reproducible shuffle, or labels rewritten into a shared vocabulary — often several of these at once.
Copying the data to do this is wasteful and loses provenance. View instead wraps a
dataset and applies an ordered pipeline of Operation steps lazily, producing a
new dataset-shaped object without modifying or duplicating the original.
When to use¶
Use a View whenever you need a subset or a rewritten version of a dataset:
building train/test splits, focusing an experiment on a few classes, balancing an
imbalanced dataset, capping a dataset for a quick run, or conforming labels before merging
datasets. A view is accepted anywhere a dataset is — pass it straight to
Embeddings, Metadata, or any evaluator.
What you will need¶
Any dataset implementing
AnnotatedDataset.A Python environment with the following packages installed:
dataevalmaite-datasets
Getting started¶
Let’s import the required libraries needed to set up a minimal working example
import numpy as np
from maite_datasets.image_classification import MNIST
from dataeval.config import set_seed
from dataeval.data import ClassFilter, Limit, Operation, Relabel, Shuffle, View
set_seed(0) # For reproducibility
A dataset to work with¶
We use the MNIST test split (10,000 handwritten digits). Any
AnnotatedDataset works the same way.
mnist = MNIST("./data", image_set="test", download=True)
print("source dataset size:", len(mnist))
print("source vocabulary:", mnist.metadata.get("index2label"))
source dataset size: 10000
source vocabulary: {0: 'zero', 1: 'one', 2: 'two', 3: 'three', 4: 'four', 5: 'five', 6: 'six', 7: 'seven', 8: 'eight', 9: 'nine'}
The basics¶
A View takes a dataset and a list of operations. It behaves like a dataset —
len(), indexing, and iteration all work — but nothing is copied, and the source is
untouched.
view = View(mnist, operations=[ClassFilter([0, 1]), Limit(500)])
print("view size:", len(view))
print("source is unchanged:", len(mnist))
print("first 10 source indices behind the view:", view.resolve_indices()[:10])
view size: 500
source is unchanged: 10000
first 10 source indices behind the view: [2, 3, 5, 10, 13, 14, 25, 28, 29, 31]
Order matters¶
Operations run in the order you list them, and each one sees the result of the previous. This makes a pipeline read top-to-bottom, but it also means order is meaningful — the same two operations in opposite orders answer different questions.
limit_then_filter = View(mnist, operations=[Limit(500), ClassFilter([0, 1])])
filter_then_limit = View(mnist, operations=[ClassFilter([0, 1]), Limit(500)])
print("Limit(500) -> ClassFilter([0, 1]):", len(limit_then_filter), "(0s and 1s *within* the first 500)")
print("ClassFilter([0, 1]) -> Limit(500):", len(filter_then_limit), "(the first 500 0s and 1s)")
Limit(500) -> ClassFilter([0, 1]): 109 (0s and 1s *within* the first 500)
ClassFilter([0, 1]) -> Limit(500): 500 (the first 500 0s and 1s)
Because each operation applies to the current selection, an operation can appear more than once. This pipeline takes a reproducible random sample from a specific window of the dataset — something a single pass could not express.
windowed = View(mnist, operations=[Limit(2000), Shuffle(seed=0), Limit(10)])
print("10 random items drawn from the first 2,000:", windowed.resolve_indices())
10 random items drawn from the first 2,000: [1946, 1236, 1380, 1949, 1633, 474, 1815, 935, 470, 1626]
Operations see the result of earlier operations¶
Operations that inspect the data read it through the operations before them. That means a filter placed after a relabeling step matches against the new vocabulary, not the original one.
Here we collapse the ten digits into even/odd with Relabel, then filter on the
relabeled classes.
digit_names = mnist.metadata.get("index2label", {})
EVEN_ODD = ("even", "odd")
parity = {name: (EVEN_ODD[digit % 2]) for digit, name in digit_names.items()}
def source_digits(v: View) -> list[int]:
"""The original MNIST digits sitting behind a view's items."""
return sorted({int(np.argmax(mnist[i][1])) for i in v.resolve_indices()})
relabel_first = View(mnist, operations=[Relabel(parity, EVEN_ODD), ClassFilter([0])])
filter_first = View(mnist, operations=[ClassFilter([0]), Relabel(parity, EVEN_ODD)])
print("vocabulary after Relabel:", relabel_first.metadata.get("index2label"))
print()
print("Relabel -> ClassFilter([0]):", len(relabel_first), "images, source digits", source_digits(relabel_first))
print("ClassFilter([0]) -> Relabel:", len(filter_first), "images, source digits", source_digits(filter_first))
vocabulary after Relabel: {0: 'even', 1: 'odd'}
Relabel -> ClassFilter([0]): 4926 images, source digits [0, 2, 4, 6, 8]
ClassFilter([0]) -> Relabel: 980 images, source digits [0]
The same two operations in the opposite order give very different results. With Relabel
first, ClassFilter([0]) reads each datum through it, so class 0 means the new
vocabulary’s even — every even digit is kept. With ClassFilter first, class 0 still
means the source digit zero, so only zeros survive.
Writing your own Operation¶
Subclass Operation and implement apply(view). Inside apply you have three
levers, and you may use any combination of them:
lever |
how |
use it for |
|---|---|---|
cardinality / order |
set |
filtering, reordering, limiting, resampling |
content |
|
rewriting each datum |
metadata |
override |
rewriting dataset-level metadata |
view.selection is the list of surviving source indices. view.map(fn) registers a
per-datum transform that is applied lazily on access; pass where={...} to target only
specific source indices. If your operation reads each datum’s target, set the class
attribute requires so the view can validate the dataset shape up front.
class KeepEvery(Operation):
"""Cardinality: keep every nth item of the current selection."""
def __init__(self, n: int) -> None:
self.n = n
def apply(self, view: View) -> None:
view.selection = view.selection[:: self.n]
class Invert(Operation):
"""Content: invert pixel intensities. Registers a transform, reads no data here."""
def apply(self, view: View) -> None:
view.map(lambda datum: (255 - np.asarray(datum[0]), datum[1], datum[2]))
custom = View(mnist, operations=[Limit(100), KeepEvery(10), Invert()])
print("custom view size:", len(custom))
print("source indices:", custom.resolve_indices())
print(
"a background pixel — source:",
int(np.asarray(mnist[0][0])[0, 0, 0]),
"view:",
int(np.asarray(custom[0][0])[0, 0, 0]),
)
custom view size: 10
source indices: [0, 10, 20, 30, 40, 50, 60, 70, 80, 90]
a background pixel — source: 0 view: 255
Note that Invert never touched the data while building the view — it only registered a
transform, which runs when an item is read. Operations that filter must read the data to
decide what to keep, so they scan once at construction; pure transforms stay free until
access.
Built-in operations¶
operation |
what it does |
|---|---|
keep datums containing one of the given classes (and mask other detections for object detection) |
|
resample to a balanced class distribution, optionally to a target size |
|
cap to the first n of the current selection |
|
randomize order, optionally seeded |
|
reverse the current order |
|
keep the given source indices |
|
rewrite labels into a target vocabulary, dropping out-of-vocabulary datums |
Tracking what a view did¶
A view records its own lineage, which is useful for reproducibility and for writing sidecar metadata alongside derived datasets.
nested = View(View(mnist, [ClassFilter([0, 1])]), [Limit(25)])
print("operations, innermost first:", nested.operation_groups)
print("original dataset:", type(nested.root_dataset).__name__)
print("source indices behind the first 5 items:", nested.resolve_indices()[:5])
operations, innermost first: [[ClassFilter(classes=[0, 1], filter_detections=True)], [Limit(size=25)]]
original dataset: MNIST
source indices behind the first 5 items: [0, 1, 2, 3, 4]
Summary¶
Viewwraps a dataset with an ordered pipeline ofOperationsteps, lazily and without copying.Operations run in the order given, each seeing the result of the previous — so order is meaningful, and an operation may appear more than once.
Operations that inspect data read it through earlier operations.
Write your own by subclassing
Operationand implementingapply(view).