Performance
PDEForge uses NumPy and SciPy for broad compatibility. For most research use cases, this is sufficient.
Typical Performance
| Dataset Size | Spectral Models | FEniCSx Models |
|---|---|---|
| 100 samples | seconds | minutes |
| 1,000 samples | minutes | tens of minutes |
| 10,000 samples | hours | not recommended |
Recommendations
Generate Once, Reuse
For training datasets, generate once and save:
dataset = generate_dataset("burgers_1d", n_samples=10000, ...)
dataset.save("./burgers_training_data")
# Later
dataset = load_dataset("./burgers_training_data")
Start Small
Begin with small datasets for development:
# Development
small = generate_dataset(model, n_samples=100, resolution={"x": 64})
# Production
large = generate_dataset(model, n_samples=10000, resolution={"x": 256})
Reduce Resolution for Exploration
When exploring parameters:
# Quick exploration at low resolution
explore_parameter(..., resolution={"x": 64, "y": 64})
# Final dataset at target resolution
generate_dataset(..., resolution={"x": 256, "y": 256})
FEniCSx Models
For cylinder flow and other FEniCSx models:
- Keep sample counts modest (50-200 samples)
- Use coarser output grids when possible
- The mesh resolution (
_mesh_resolution) affects solve time significantly
Parallel Generation
Some speedup is possible via parallel workers:
This parallelizes sample generation. Speedup depends on solver overhead and system resources.
Memory Considerations
Large 2D datasets can consume significant memory:
For very large datasets:
- Generate in batches
- Save each batch to disk
- Concatenate when loading
Future: GPU Acceleration
A JAX-based backend for spectral models is planned for future releases. This would enable GPU-accelerated data generation for models like Burgers and Darcy.