01 · SLIDE 1 · OPENING
Title & Introduction ~25 sec
Good afternoon, respected conference chair, distinguished researchers, colleagues, and participants. It is a pleasure to be here at ICORIS 2026. My name is Dedi Julyan Sukawanto from Dipa Makassar University, Indonesia. Today, I am pleased to present our paper entitled, “Characterizing Request Patterns of Web Crawlers on a Regional News Website Using Browser Fingerprinting.” Our Paper ID is six hundred sixty-two (662).
ICORIS = International Conference on Cybernetics and Intelligent Systems → “eye-COR-iss” · ID = Identification → “eye-dee”
02 · SLIDE 2 · BACKGROUND & MOTIVATION
Why This Matters ~25 sec
AI-powered web crawlers increasingly access content-rich websites. Regional news sites are sensitive to server resource use and analytics distortion. Browser fingerprinting provides device-level and rendering-level signals beyond standard logs.
The dataset contains three thousand two hundred forty-one (3,241) raw fingerprint records and two thousand six hundred nine (2,609) valid records. Bot traffic represents fifty-two point five percent (52.5%), while human traffic represents forty-seven point five percent (47.5%).
AI = Artificial Intelligence → “A-I”
03 · SLIDE 3 · RESEARCH GAP
Research Gap ~20 sec
Previous crawler-detection studies often emphasize server-side logs or classification accuracy. Browser-fingerprinting studies commonly focus on uniqueness, scripts, or headless browser detection. However, limited work characterizes AI crawlers, search engine crawlers, and human users together, especially on regional news websites.
Our study extends browser fingerprinting toward an interpretable characterization of crawler and human traffic.
04 · SLIDE 4 · OBJECTIVES & CONTRIBUTIONS
Objectives & Contributions ~20 sec
Our objectives are to characterize traffic composition, quantify group differences using browser fingerprint and request-pattern features, and evaluate two proposed metrics: FUR and DAI. Overall, we contribute six features, two proposed metrics, empirical regional evidence, and a preprocessing pipeline.
FUR = Fingerprint Uniqueness Ratio → “F-U-R” · DAI = Device Anomaly Index → “D-A-I”
05 · SLIDE 5 · DATASET & INITIAL GROUPING
Dataset & Initial Grouping ~25 sec
The collection period was from November first to December thirtieth, twenty twenty-five. Records include timestamp, source IP, User Agent, page URL, page title, and fingerprint components such as Canvas, WebGL, CPU concurrency, device memory, touchscreen, plugins, fonts, locale, and timezone.
The groups are: AI Crawlers — one thousand one hundred four (1,104), or forty-two point three percent (42.3%); Search Engine Crawlers — two hundred sixty-five (265), or ten point two percent (10.2%); Human Users — one thousand two hundred forty (1,240), or forty-seven point five percent (47.5%).
IP = Internet Protocol → “I-P” · URL = Uniform Resource Locator → “U-R-L” · WebGL = Web Graphics Library → “web G-L” · CPU = Central Processing Unit → “C-P-U”
06 · SLIDE 6 · PROPOSED METHODOLOGY
Research Workflow ~15 sec
The proposed methodology consists of data collection, preprocessing, feature extraction, comparative analysis, and crawler characterization. The workflow separates initial traffic grouping from fingerprint-based characterization.
07 · SLIDE 7 · PREPROCESSING & FEATURES
Preprocessing & Features ~25 sec
We use six features: CPU concurrency, device memory, inter-request time, Fingerprint Uniqueness Ratio, Device Anomaly Index, and Page Type Distribution. FUR captures rendering-hash homogeneity, while DAI captures hardware anomalies from browser-exposed CPU and memory values. Because the feature distributions are skewed, nonparametric tests were used.
Preprocessing removed six hundred thirty-two (632) duplicate records, normalized two thousand three hundred eighty-eight (2,388) language values, fixed one hundred twelve (112) memory values, and flagged one thousand two hundred eighty-nine (1,289) CPU anomalies.
IRT = Inter-Request Time → “I-R-T” · PTD = Page Type Distribution → “P-T-D”
08 · SLIDE 8 · TRAFFIC COMPOSITION
Traffic Composition ~15 sec
AI Crawlers account for forty-two point three percent (42.3%) of valid records and appear about four times more frequently than Search Engine Crawlers. Human Users account for forty-seven point five percent (47.5%).
The figure also shows the distribution of bot names across the traffic groups.
09 · SLIDE 9 · FEATURE COMPARISON
Comparison Across Groups ~25 sec
The feature comparison shows clear differences between the three groups. Reported CPU concurrency is much higher for crawlers: one hundred twenty-two point zero (122.0) for AI Crawlers, one hundred thirty-five point zero (135.0) for Search Crawlers, and seven point seven (7.7) for Human Users.
Device memory is eight point zero gigabytes (8.0 GB) for both crawler groups and five point four gigabytes (5.4 GB) for Human Users. Average inter-request time is four thousand three hundred sixty-six point five seconds (4,366.5 s) for AI, one thousand eighty-nine point eight seconds (1,089.8 s) for Search, and two thousand four hundred seventy-two point two seconds (2,472.2 s) for Humans.
FUR = Fingerprint Uniqueness Ratio → “F-U-R” · IRT = Inter-Request Time → “I-R-T”
10 · SLIDE 10 · HARDWARE FINGERPRINT FINDINGS
Hardware Fingerprint Findings ~20 sec
AI and Search Engine Crawlers show server-grade or virtualized hardware profiles. The Device Anomaly Index flagged ninety-four point two percent (94.2%) of AI Crawler requests, compared with zero percent (0%) of Human User requests. Touchscreen prevalence is high for Human Users and low or absent for crawler groups.
DAI = Device Anomaly Index → “D-A-I” · CPU = Central Processing Unit → “C-P-U”
11 · SLIDE 11 · FINGERPRINT UNIQUENESS
Fingerprint Uniqueness ~20 sec
Bot traffic produced only ten (10) unique Canvas hashes, while Human Users produced one hundred seventy-nine (179). AI Crawler Canvas FUR was zero point zero zero five four (0.0054), compared with zero point one four four four (0.1444) for Human Users.
Therefore, Human User Canvas FUR was approximately twenty-six point seven times higher (26.7 times). This shows that FUR provides an interpretable measure of rendering-environment concentration.
FUR = Fingerprint Uniqueness Ratio → “F-U-R”
12 · SLIDE 12 · BEHAVIORAL PATTERNS
Behavioral Patterns ~20 sec
Human User requests concentrate on article pages. AI Crawlers access tag, category, search, and navigational pages more frequently. Request volume also increases during afternoon and evening activity windows.
13 · SLIDE 13 · KEY TAKEAWAYS & LIMITATIONS
Key Takeaways & Limitations ~25 sec
Bot traffic is dominated by English, followed by Chinese and Russian, while Human User traffic is predominantly Indonesian, reflecting regional readership.
In conclusion, FUR captures rendering homogeneity, while DAI captures non-consumer hardware profiles. Both should be used as complementary characterization metrics, not standalone detection rules. The main limitations are User Agent spoofing, proxy-mediated requests, browser privacy protections, and the single-site scope.
Future work can add ASN, reverse DNS, IP reputation, and independently verified labels.
14 · SLIDE 14 · THANK YOU
Closing ~10 sec
Thank you very much for your attention. That concludes my presentation. I would be happy to answer any questions.
Questions & Discussion → “Questions and discussion.”