01 — Doctoral research · BINUS University

TANDAN

Technology for Affordable Near-field Drone-based Assessment of Oil Palm Plantations

Affordable eyes in the sky for oil palm plantations.

5tree conditions detected
17 FPSsustained engine throughput on a Jetson Nano
0.9694mAP@0.5:0.95 at 768 px (0.9648 at the deployed 704 px)
TANDAN drone in flight over an oil palm plantation
The TANDAN drone flying between oil palm rows during field testing.

02/14The problem

Spotting a struggling palm among thousands

Indonesia is the world's largest palm oil producer. Oil palm contributes roughly 17% of Indonesia's agricultural GDP and employs up to 7.8 million people across its value chain (Purnomo et al., 2020), and palm oil supplies about 40% of global vegetable oil demand (Qaim et al., 2020). Indonesia's palm oil exports reached USD 22.97 billion in 2020 (Xin et al., 2022).

Keeping plantations productive means spotting problem trees early: dead palms, yellowing crowns, poorly maintained trees and stunted growth. Diseases such as basal stem rot (Ganoderma boninense) often become visible only at a late stage. Walking thousands of hectares to inspect trees by eye is slow, costly and inconsistent. Commercial drone-analytics services exist, but they are expensive and usually depend on cloud processing that remote plantations cannot rely on.

~17%of Indonesia's agricultural GDP (Purnomo et al., 2020)
7.8 Mpeople employed across the value chain, at most (Purnomo et al., 2020)
~40%of global vegetable oil demand (Qaim et al., 2020)
USD 22.97 BIndonesian palm oil exports in 2020 (Xin et al., 2022)
“The question was not whether AI can recognize a struggling palm — it was whether it can do so on hardware a plantation can actually afford.”

03/14What TANDAN is

An end-to-end system in three parts

01

A custom drone

Built on an S500 quadcopter frame, carrying an NVIDIA Jetson Nano as its onboard computer.

02

A lightweight detection model

YOLO11n, recognizing five tree conditions, optimized to run on low-power edge hardware.

03

TANDAN-V1

Monitoring software that tracks each tree, counts unique palms, computes a plantation health score and exports survey reports.

Two operating modes

ModeHow it worksStatus
Onboard edge modeCamera on the drone → Jetson Nano runs the optimized TensorRT engine on the aircraftEngine validated on the Jetson Nano; in-flight integration in progress
Ground-station mode (TANDAN-V1)Drone video streamed over RTMP → laptop runs detection, tracking and the live dashboardOperational; used in field testing

04/14The drone

An open, custom-built airframe

TANDAN uses a custom-built UAV rather than a closed commercial drone, so that the flight controller, onboard computer and camera can be integrated openly for research.

The custom S500-based TANDAN drone
    FIG. 4.1The custom S500-based TANDAN drone. Select a numbered marker to label a component.

    The build uses an open ArduPilot flight stack and off-the-shelf hobby-grade components. A closed agricultural flight controller (DJI N3-AG) was evaluated and not adopted, because its ecosystem does not expose the open interfaces that onboard AI integration needs.

    Specifications
    FrameS500 quadcopter frame, 480–500 mm class
    Flight controllerAPM 2.6 (ArduPilot Mega)
    FirmwareArduPilot, configured with Mission Planner
    Motors2212 920 KV brushless (×4)
    ESCs20 A BLHeli-S (×4)
    Propellers1045 (10 × 4.5 in)
    BatteryLiPo 3S 2200 mAh
    Radio transmitterFrSky Taranis X9D Plus
    RF linkRadioMaster Ranger Micro 2.4 GHz ExpressLRS (ELRS) external TX module with receiver
    Onboard computerNVIDIA Jetson Nano 4 GB (JetPack 4.6.6, CUDA 10.2, TensorRT 8.2.1)
    CameraArducam IMX219 CSI camera (8 MP; sensor modes up to 3264 × 2464 @ 21 fps and 1280 × 720 @ 120 fps)

    Measured camera capture through the Jetson's hardware image pipeline: about 28.8 FPS at 1280 × 720 — well above the 15 FPS real-time threshold, so capture is not the bottleneck.

    05/14The detection model

    Five conditions, one very uneven dataset

    The model was trained on MOPAD (Multi-class Oil PAlm Detection dataset; Zheng et al., 2021), Site 2 subset.

    1,803training images
    500validation images
    1024 × 1024pixels per image
    151,583annotated trees
    Five condition classes
    Class (EN)Class (ID)MeaningTraining instances
    healthynormalHealthy mature palm89,958
    smallishkecilSmall or stunted growth26,644
    yellowishmenguningYellowing crown, possible nutrient or disease stress1,473
    mismanagedtidak terawatPoorly maintained, overgrown surroundings431
    deadmatiDead palm231

    The data is heavily imbalanced. The mismanaged class is both rare and hard, because it is recognized mainly from fine texture around the crown.

    389 : 1healthy palms for every dead palm in training
    Full class distribution
    SetdeadhealthymismanagedsmallishyellowishTotal
    Training23189,95843126,6441,473118,737
    Validation6124,884957,39541132,846
    Total292114,84252634,0391,884151,583

    06/14The research journey

    Nine ideas, then one change that mattered

    We tested nine ways to make the model smarter. None was adopted. Then we changed one thing — the image resolution.

    Step 1

    Choosing a base model

    Four nano-sized YOLO models were compared in the same screening run.

    ModelmAP@0.5:0.95Params (M)GFLOPsFPS on JetsonReal-time?
    YOLOv8n0.9505 ± 0.00703.0068.121.8Yes
    YOLO11n0.9578 ± 0.00232.5836.3220.2Yes — selected
    YOLO12n0.9696 ± 0.00262.5586.329.8No
    YOLO26n0.9394 ± 0.0074——export failedNo

    YOLO12n was the most accurate but ran at only 9.8 FPS on the Jetson Nano, because its attention mechanism relies on hardware (FlashAttention / Tensor Cores) that the Nano's Maxwell GPU lacks. YOLO26n could not be exported to TensorRT. YOLO11n offered the best balance of accuracy, size and stability.

    In the main experiments that followed, the YOLO11n baseline at 640 px measured 0.9619 ± 0.0023 — this is the reference used for every comparison below.

    Step 2

    Nine modifications, none adopted

    Each was tested in isolation over the same baseline and judged against a statistical noise floor: the standard deviation of the baseline across three training seeds (42, 123, 456), ≈ 0.0023.

    Group A — tested on YOLO12n, the accuracy leader
    ModificationmAP@0.5:0.95ChangeCostVerdict
    YOLO12n baseline0.9696 ± 0.0026—2.558 M params, 6.32 GFLOPsreference
    SPD-Conv (P3–P5)0.9692 ± 0.0048−0.00043.995 M params, 9.86 GFLOPs (+56%)neutral, expensive
    Varifocal Loss0.9756 (seed 42)+0.0060unchangednot adopted — background false positives rose 2.2× (364 → 809)
    A2C2fRep (attention removed)0.9565 ± 0.0031−0.01312.927 M params, 6.79 GFLOPsdeclined

    SPD-Conv is a small-object technique, but oil palm crowns are medium-sized objects (about 62 px), so it added 56% compute for no gain.

    Group B — tested on YOLO11n (baseline 0.9619 ± 0.0023; seed-42 reference 0.9593 for single-seed rows)
    ModificationPart changedmAP@0.5:0.95Verdict
    Class-balanced focal loss (Cui et al., 2019)Loss0.9597 ± 0.0021neutral
    SCS-Head (shared-convolution head)Head0.9156failed
    SimAM parameter-free attention (Yang et al., 2021)Backbone0.9501failed
    RepC3k2 re-parameterization (after Ding et al., 2021)Backbone0.9509declined
    BiFPN weighted feature fusion (Tan et al., 2020)Neck0.9586neutral
    Network Slimming structured pruning (Liu et al., 2017)Whole networkchannels did not polarize — haltedfailed
    Pruning failure evidence
    BN scale statisticBeforeAfterNeeded
    Mean absolute value1.10451.0644≪ 1.0
    Channels < 0.010.013%0.013%> 20%
    Channels < 0.100.026%0.026%> 30%

    The class-balanced loss shifted where false positives went (background → healthy dropped from 66.3% to 52.6%) but did not remove them — they moved to smallish (31.0% → 42.1%).

    Step 3

    What worked: resolution

    Input resolutionmAP@0.5:0.95GFLOPsAP mismanagedSeeds
    640 px0.9619 ± 0.00236.320.8943
    704 px0.96487.640.9081
    768 px0.9694 ± 0.00269.10.9263

    Raising resolution from 640 to 768 px improved mAP by 0.0075 — more than three times the noise floor — and brought YOLO11n level with YOLO12n (0.9696) while remaining deployable. The hardest class, mismanaged, gained over three points.

    Try it — input resolution

    640704768
    mAP@0.5:0.95—
    AP mismanaged—
    GFLOPs—
    Jetson FPS (TensorRT FP16)—

    Step 4

    Why it worked

    Per-class accuracy across eight configurations. Spread of each class across configurations: healthy 0.021, dead 0.043, yellowish 0.048, smallish 0.050, mismanaged 0.125 — two to six times wider than any other class. Ranking configurations by mismanaged accuracy almost reproduces the ranking by overall mAP; the other four classes were already saturated.

    Hover or focus a point to see its configuration. Gold = resolution 768.

    AP per class
    Configurationdeadhealthymismanagedsmallishyellowish

    Key insight

    The bottleneck was pixels, not model capacity. Fine crown texture disappears when images are downsampled, and no module can recover information that is not in the input. The nine rejected modifications are not wasted effort — they are the evidence that ruled every other explanation out.

    Error diagnosis. Most errors were not confusion between classes but background mistaken for trees — about 66% of background false positives were labeled healthy and 31% smallish.

    Overall accuracy tracks the mismanaged class
    FIG. 6.1Overall mAP against AP of the mismanaged class across configurations.
    Accuracy versus compute trade-off
    FIG. 6.2Accuracy versus compute on the Jetson Nano, marked by real-time feasibility.

    07/14Edge deployment on Jetson Nano

    Nine engines, one that clears the bar

    1. PyTorch.pt
    2. ONNXopset 12
    3. TensorRT enginebuilt with trtexec
    4. NVIDIA Jetson Nano 4 GBMaxwell GPU, compute capability 5.3

    Throughput (queries per second) · three resolutions × three precisions

    11.220.2 qps
    1. INT8 gives no speedup on this GPU — within 0.3% of FP32 at every resolution. The Maxwell GPU lacks the DP4A instruction that accelerates 8-bit math (available only from compute capability 6.1). Measured, not assumed.
    2. FP16 gives a consistent ~30% speedup at every resolution.
    3. 768 px fails the 15 FPS real-time threshold (14.52) despite being most accurate; 704 px passes with a 13% margin (17.02). 704 px FP16 is the deployed configuration.
    Engine throughput against the real-time threshold
    FIG. 7.1Engine throughput against the 15 FPS real-time threshold.

    FP16 accuracy check

    Measured on the Jetson, using PyTorch half precision as a numerically equivalent proxy because TensorRT's Python bindings are unavailable for Python 3.8 on JetPack 4.6.

    ClassFP32FP16Δ
    dead0.9790.981+0.002
    healthy0.9930.9930.000
    mismanaged0.9310.927−0.004
    smallish0.9770.971−0.006
    yellowish0.9880.986−0.002
    all0.9740.971−0.003

    Thermal stress test

    313 s of continuous inference.

    MetricIdleUnder load
    GPU temperature23.1 °C21.5 → 30.5 °C, plateau, max 35.0 °C
    CPU temperature21.4 °C21.5 → 31.5 °C, max 36.5 °C
    GPU utilization~0%stable 89–98%
    Throughput—17.0025 qps sustained vs 17.02 instantaneous (Δ 0.1%)

    No thermal throttling: temperatures plateaued more than 60 °C below the ~97 °C throttling point, and sustained throughput matched the instantaneous value.

    Final configuration

    Model
    YOLO11n
    Input
    704 px
    Engine
    TensorRT FP16
    Throughput
    17 FPS sustained
    Parameters
    2,583,127
    Output
    1 × 9 × 10,164

    08/14TANDAN-V1 monitoring software

    From detections to a survey a manager can use

    TANDAN-V1 turns raw detections into survey information a plantation manager can use.

    1. Drone flight
    2. Videolive RTMP stream via mediamtx, or recorded file
    3. YOLO detection
    4. ByteTrack multi-object tracking
    5. Per-tree classificationby majority vote
    6. Live dashboard
    7. Automatic session report
    VID. 8.1Demo of TANDAN-V1 tracking palms: annotated video on the left, analytics panel on the right. Figures visible in the recording belong to a demo session and are not reported results.
    • Split-screen dashboard: annotated video (about 65% of the screen) and analytics panel (about 35%).
    • Box labels show track ID · class · confidence.
    • Unique palms counting: each tree keeps a stable track ID; when it leaves view for 15 frames it is retired and its final class is decided by majority vote across all its observations. The dashboard shows both detections (every frame) and unique palms (each tree once) — unique counts are what matter for inventory.
    • Plantation health score (0–100) with a traffic-light gauge: green ≥ 75, yellow 50–75, red < 50.
    • Alerts: a flashing red border appears when a dead or mismanaged palm is in view.
    • Live FPS and detection sparklines, confidence histogram, per-class counts.
    • Exports per session: annotated MP4, CSV (frame, timestamp, elapsed seconds, track ID, class, confidence, box coordinates, inference time, FPS) and a JSON summary.
    • Post-flight analytics page: loads a session CSV and shows class distribution, timeline, confidence and FPS charts, and a table of at-risk palms.
    TANDAN-V1 live monitoring dashboard shown on a tablet in the field
    FIG. 8.2TANDAN-V1 live monitoring dashboard on a tablet.
    Post-flight analytics of a survey session
    FIG. 8.3Post-flight analytics of a demo session.

    Health score

    good  = healthy + 0.7 × smallish
    bad   = dead + 0.8 × mismanaged + 0.4 × yellowish
    score = good / (good + bad) × 100
            (computed on unique palms)

    Adjust the five class counts to see how the score and gauge colour respond.

    —
    —

    09/14Field testing

    In the plantation

    TANDAN was tested at an oil palm plantation in Kabupaten Tanjung Jabung Barat, Jambi, Indonesia. Aerial imagery was processed into orthomosaics with Agisoft Metashape, and the detector — trained on data from a different region — was applied to the new site's imagery.

    Research team during field testing, with the TANDAN drone and ground-station laptop
    Research team during field testing.
    Mission planning on a handheld controller at the plantation
    Planning an aerial survey at the test site.
    Drone take-off at the test site.

    10/14Results at a glance

    What the study established

    0.9694

    mAP@0.5:0.95 at 768 px — level with a newer, heavier model

    0.9648

    mAP at the deployed 704 px

    +0.0075

    mAP from resolution alone (3× the noise floor)

    9

    modifications tested, none adopted

    17 FPS

    sustained on a Jetson Nano, 13% above the real-time threshold

    ~30%

    speedup from FP16; none from INT8

    −0.003

    mAP cost of FP16

    60 °C+

    thermal headroom under continuous load

    2.58 M

    parameters

    11/14Network architecture

    Stock YOLO11n, deliberately

    The deployed network is stock YOLO11n. The contribution is the resolution choice and the deployment pipeline, not a new architecture.

    • BackboneLayers 0–10: Conv → Conv → C3k2 → Conv → C3k2 → Conv → C3k2 → Conv → C3k2 → SPPF → C2PSA; channel widths 16 → 256.
    • NeckLayers 11–22: PAN-FPN with upsampling and concatenation skip connections from layers 4, 6 and 10.
    • HeadLayer 23: detection at three scales — P3/8 (88 × 88 grid at 704 px), P4/16 (44 × 44), P5/32 (22 × 22) — decoded to 10,164 candidate boxes (4 box coordinates + 5 class scores).
    • Size2,583,127 parameters · 101 fused layers · 6.3 GFLOPs at 640 px
    Deployed YOLO11n network architecture
    FIG. 11.1Deployed YOLO11n network architecture. Select to enlarge.
    Methodology from dataset to deployment-ready engine
    FIG. 11.2Methodology from dataset to deployment-ready engine. Select to enlarge.

    12/14Status

    Where the project stands

    1. Dataset preparation and model trainingDone
    2. Nine-modification study and resolution studyDone
    3. Jetson Nano engine benchmark, FP16 validation, thermal testDone
    4. TANDAN-V1 monitoring softwareDone
    5. Custom S500 drone buildDone
    6. Field testingDone
    7. Live onboard camera pipeline on the JetsonIn progress
    8. In-flight onboard detectionPlanned

    13/14Team

    People

    Researcher
    Muhammad Alfhi Saputra, S.Kom., M.Kom.
    Promotor
    Prof. Dr. Ir. Widodo Budiharto, S.Si., M.Kom., IPM., SMIEEE
    Co-promotors
    Dr. Ir. Haryono Soeparno, M.Sc.
    Dr. Ir. Yulyani Arifin, S.Kom., M.M., IPP
    Institution
    BINUS University
    Dataset credit
    MOPAD dataset by Zheng et al. (2021)

    Dissertation

    Development of an Efficient Real-Time UAV-Based Detection System for Oil Palm Condition Monitoring

    Pengembangan Sistem Deteksi Berbasis Real-Time-UAV yang Efisien untuk Monitoring Kondisi Kelapa Sawit

    Doctoral dissertation, Doctor of Computer Science.

    References

    Sources

    1. Zheng, J. et al. (2021). Growing status observation for oil palm trees using UAV images. ISPRS J. Photogramm. Remote Sens. 173, 95–121. (MOPAD dataset)
    2. Purnomo, H. et al. (2020). Reconciling oil palm economic development and environmental conservation in Indonesia. For. Policy Econ. 111, 102089.
    3. Qaim, M. et al. (2020). Environmental, economic, and social consequences of the oil palm boom. Annu. Rev. Resour. Econ. 12, 321–344.
    4. Xin, Y., Sun, L., Hansen, M.C. (2022). Oil palm reconciliation in Indonesia. J. Clean. Prod. 380, 135087.
    5. Cui, Y. et al. (2019). Class-balanced loss based on effective number of samples. CVPR.
    6. Yang, L. et al. (2021). SimAM: a simple, parameter-free attention module. ICML.
    7. Tan, M., Pang, R., Le, Q.V. (2020). EfficientDet: scalable and efficient object detection. CVPR.
    8. Ding, X. et al. (2021). RepVGG: making VGG-style ConvNets great again. CVPR.
    9. Liu, Z. et al. (2017). Learning efficient convolutional networks through network slimming. ICCV.
    10. Jocher, G., Qiu, J. (2024). Ultralytics YOLO11. github.com/ultralytics/ultralytics