NirmaanAI · Testing & Validation Report

NirmaanAI — Testing & Validation Report (PoC)

Rodic InfraAI Innovation Challenge 2026 · BRS v2 §24–29, §43.3 · computed 29 Sep 2026.

Live application: https://nirmaanai.tecell.in · Live evidence page: https://nirmaanai.tecell.in/evidence · Demo video: https://youtu.be/XJhZFgefzBI · Contact: [email protected]

Every figure below is produced by the pipeline and can be re-derived from the live GET https://nirmaanai.tecell.in/api/metrics endpoint. Nothing is hand-entered.

1. Dataset-level testing (BRS §24.1, §28)

Check Result
Source I-BLEND (IIIT Delhi), CC0, 16 energy CSVs + occupancy + calendar (see docs/dataset.md)
Records ingested 25.1 million 1-minute meter records → 112,049,266 individual measurements (13 entities × up to 5 parameters)
Time coverage 2013-08-10 → 2017-12-31 (IST); overall power coverage 86.1 %
Timestamp conversion Passed: 1376073000 → 2013-08-10 00:00 +05:30; all timestamps within range
Duplicate timestamps 0 in every file
Out-of-range values 7,194 flagged (retained with flag); 104 suspect_zero (0 V with 0 W)
Source consistency Per-entity power vs all_buildings_power.csv / all_transformer_power.csv: 100 % of joined rows within 0.5 W
Availability file agreement data_present_status_*.csv vs our own presence: 100 %
Largest gap Transformer 1: 679,510 min (~472 days); reported as missing, never as anomalies

Physical validity ranges are applied per entity type (building and transformer meters are checked against different electrical ranges).

2. Anomaly detection (BRS §24.2, §25)

Method (overview). Each entity's expected behaviour is learned from its own history, taking time and calendar context into account. Statistical and machine-learning detectors then evaluate deviations in consumption and in electrical supply-quality parameters. Periods with insufficient data are never treated as anomalies (BRS §7.4).

Ground truth. The dataset has no anomaly labels, so detection is measured on synthetic anomalies injected into copies of clean 3-day windows from the last 30 % of each entity's timeline (40 windows × 10 eligible entities, seeds 1–3). Injected data never enters the live database. Two magnitude tiers are reported, and the detectors were not tuned to them.

Anomaly type Subtle magnitude Significant magnitude
spike power ×2.5, 1 bucket same
sustained increase / decrease +40 % / −50 %, 8 buckets same
voltage deviation −10 %, 4 buckets same
frequency deviation +0.3 Hz, 3 buckets set to 50.7 Hz, 3 buckets
power-factor drop |PF| → 0.70, 6 buckets |PF| → 0.60, 6 buckets
current spike (power unchanged) ×2, 4 buckets same
missing-data sequence 8 buckets invalid (expect no event) same

Results (mean over 3 seeds, 400 injections per tier).

Type P (sig.) R (sig.) F1 (sig.) F1 (subtle) FPR (sig.) Latency (buckets)
spike 1.00 0.76 0.86 0.86 0.013 0.0
sustained increase 1.00 0.59 0.74 0.74 0.011 1.1
sustained decrease 1.00 0.74 0.85 0.84 0.011 0.8
voltage deviation 1.00 0.88 0.93 0.93 0.005 0.0
frequency deviation 1.00 0.94 0.97 0.04 0.004 0.0
power-factor drop 0.66 0.87 0.75 0.66 0.048 0.0
current spike 0.62 0.83 0.71 0.73 0.036 0.0
Overall 0.86 0.78 0.82 0.75 0.045 0.3

Detection stability (Jaccard of detected injections when thresholds are perturbed ±10 %): 0.995. Across-seed standard deviation of F1 ≤ 0.04 per type. Missing-data sequences raised a (false) event in 4 % of cases.

Reading the results. Consumption anomalies are detected with no false detections on the injected data (precision 1.00). Recall is lower for gradual sustained increases (0.59). Subtle frequency shifts are below the alerting level by design, hence F1 0.04 in that tier. Grid-band compliance is reported as a statistic instead of as events.

Naturally occurring anomalies (real data). 35,053 events: consumption 16,711, peak 8,804, supply quality 9,538. Of the consumption events, 56.0 % start on non-working or low-activity days, against a 56.7 % base rate of such days. The calendar alone doesn't explain the detections. This is context, not ground truth.

3. Peak-demand intelligence (BRS §27)

Check Result
Peak events detected 8,804 (mean deviation vs baseline 50.2 %)
Magnitude accuracy (15-min vs true 1-min peak) mean ratio 0.90, MAPE 9.9 % (n = 8,526). 15-min averaging understates short spikes
Ranking consistency (entity ranking by daily peak) Kendall τ 0.88 between consecutive quarters; 0.97 between 15-min and 1-h resolution
Synthetic peaks (+60 %, 2–4 buckets) recall 0.50, mean timing error 2.6 buckets
High-demand identification 45.6 % of detected peak events fall in the entity's top-5 % demand buckets

4. Forecasting (BRS §26)

Chronological split: train ≤ 2016-06-30, validation 2016-07-01 → 2016-12-31, test = calendar year 2017 (never seen in training or model selection). Naive seasonal baselines and machine-learning models are compared, the best model is selected per entity and horizon on the validation period, and prediction intervals are calibrated on validation data. MAPE excludes near-zero actual values.

Entity MAE 15 min MAPE 15 min R² 15 min MAE 24 h MAPE 24 h R² 24 h
Campus 19.0 kW 5.8 % 0.90 45.4 kW 14.8 % 0.65
Transformer 1 9.1 kW 12.6 % 0.82 23.9 kW 42.1 % 0.54
Transformer 2 8.8 kW 5.6 % 0.93 26.1 kW 17.3 % 0.64
Transformer 3 4.4 kW 5.3 % 0.92 8.8 kW 10.4 % 0.73
Academic 1.5 kW 4.5 % 0.97 4.1 kW 12.1 % 0.81
Boys hostel (mains / UPS) 1.6 / 0.5 kW 8.0 / 3.2 % 0.91 / 0.98 3.8 / 1.3 kW 17.7 / 7.5 % 0.59 / 0.90
Girls hostel (mains / UPS) 0.7 / 0.2 kW 8.8 / 3.3 % 0.85 / 0.97 1.1 / 0.5 kW 14.4 / 7.1 % 0.61 / 0.87
Library 0.5 kW 6.6 % 0.97 2.0 kW 30.0 % 0.76
Mess 1.9 kW 8.6 % 0.92 3.9 kW 17.3 % 0.69
Facilities 0.5 kW 4.1 % 0.93 1.2 kW 9.4 % 0.75
Lecture 0.1 kW 35.5 % 0.90 0.5 kW 80.0 % 0.45

Against naive baselines (campus). Seasonal-naive-day MAE is 50.0 kW with R² 0.47 at every horizon. The selected model cuts the error by 62 % at 15 min (19.0 kW) and by 9 % at 24 h (45.4 kW), and raises R² from 0.47 to 0.90 and 0.65. The lecture building's high MAPE reflects its many near-zero-load periods.

5. Topology inference (BRS §44.4)

Building↔transformer links are not in the dataset, so we estimated them statistically from the measured loads. Explained variance is low (R² T1 0.14, T2 0.10, T3 0.26), and only one link is confident: Boys Hostel (mains) → T3, r = 0.46. The dashboard and assistant always label links as inferred. We don't claim a network topology.

6. Event prioritisation (BRS §19)

Events are ranked by combining how large, how long and how recurrent a deviation is with the criticality of the affected asset and parameter. Bands are calibrated on the real event population, so that Critical is reserved for the most significant sustained events (about 2 % of events). Distribution: Critical 705 · High 2,859 · Medium 13,963 · Low 17,526.

7. AI assistant (BRS §21–22)

Grounded design: the assistant answers only from the platform's own analytical results, cites the events it used, and is checked against those results before an answer is shown. Provider chain: Claude → Claude via OpenRouter → local Llama 3.1 8B (Ollama) → rule-based. Local-model benchmark (12 questions, including the 5 BRS FR-13 questions and an out-of-range date): grounding pass rate 100 % for all three local models tested. llama3.1:8b was chosen (p50 1.1 s).

8. System performance (BRS §29)

Measured on this host (32 cores, 188 GB RAM) and over the public URL through Cloudflare.

Stage Time
Ingestion (16 CSV, 1.9 GB → Parquet) 38 s
Feature generation (15-min) 2.9 s
Baselines / anomaly / supply-quality engines 130 s / 29 s / 10 s
Anomaly model training / inference 5.0 s / 21.3 s
Forecast model training (52 models) / inference 19.6 s / 14.2 s
API p50 over the internet (health, entities, overview, events, peaks, compare) ~95–140 ms
API under load (50 requests, concurrency 10) overview p50 294 ms / p95 467 ms; events p50 279 ms
Dashboard load (HTML + JS, 281 KB) 0.54 s
Assistant answer (Claude Opus 5.5 via OpenRouter, low reasoning effort, tool-grounded) ~10–11 s median (8–16 s range)
Timeseries API (one day, 15-min) ~0.1 s

9. Limitations (BRS §44)

  1. I-BLEND is one campus. Results don't generalise to DISCOM or national scale.
  2. No ground-truth anomaly or theft labels. Detection quality is measured on synthetic injections, and no event is presented as theft, failure or a fault.
  3. Transformer topology is unknown. The inferred links are weak and labelled as such.
  4. Transformer 1 has a ~472-day gap. Campus demand (= T1+T2+T3) exists only where all three report.
  5. Weather data in the collection covers 2018 only, with no overlap, so it is not used.
  6. Synthetic anomalies are clearly identified and never written to the operational data.
  7. Explanations and recommendations are generated from computed results only. They are operational guidance, not control actions.