Skip Navigation
Skip to contents

Ann Coloproctol : Annals of Coloproctology

OPEN ACCESS
SEARCH
Search

Articles

Page Path
HOME > Ann Coloproctol > Volume 42(3); 2026 > Article
Review
Anorectal benign disease
Artificial intelligence in anal fistula: mapping evidence to IDEAL stages
Vipul D Yagnik1orcid, Prema Ram Choudhary2orcid, Pankaj Garg3orcid
Annals of Coloproctology 2026;42(3):264-272.
DOI: https://doi.org/10.3393/ac.2025.01494.0213
Published online: June 25, 2026

1Department of Surgery, Banas Medical College and Research Institute, Palanpur, India

2Department of Physiology, Banas Medical College and Research Institute, Palanpur, India

3Department of Colorectal Surgery, Garg Fistula Research Institute, Panchkula, India

Correspondence to: Pankaj Garg, MBBS, MS Department of Colorectal Surgery, Garg Fistula Research Institute, 1042, Sector-15, Panchkula 134113, India Email: drgargpankaj@yahoo.com
• Received: December 4, 2025   • Revised: April 6, 2026   • Accepted: April 19, 2026

© 2026 The Korean Society of Coloproctology

This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (http://creativecommons.org/licenses/by-nc/4.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

prev next
  • 878 Views
  • 33 Download
  • 1 Web of Science
  • 2 Crossref
  • Artificial intelligence (AI) is increasingly applied in colorectal and anorectal surgery, particularly for complex conditions such as anal fistula (AF). Conventional diagnostic and therapeutic approaches remain limited by intricate anatomy, high recurrence risk, and the need to preserve continence. This narrative review, evaluates the role of AI in AF and related anorectal disorders, with evidence mapped to the IDEAL (idea, development, exploration, assessment, and long-term follow-up) framework (stages 1–4, with stage 2 subdivided into 2a and 2b). Relevant literature was identified through targeted searches of PubMed, Embase, and Scopus. Studies investigating AI applications in AF or related anorectal conditions, including imaging, surgical planning, predictive modeling, and functional assessment, were included. Evidence was categorized according to IDEAL stages, ranging from proof-of-concept to long-term quality assurance. AI demonstrates potential across 4 key domains: (1) preoperative imaging (stages 1–2b); (2) intraoperative planning and assistance (stages 1–2a); (3) predictive modeling (stages 2a–2b); and (4) broader anorectal applications, including anorectal manometry and the EndoFLIP (endoluminal functional lumen imaging probe) procedure (stages 1–2b). Feasibility studies report high diagnostic performance, particularly for magnetic resonance imaging and computed tomography–based deep learning models; however, these findings are constrained by small sample sizes, limited external validation, and challenges related to workflow integration. Overall, AI has the potential to enhance diagnostic accuracy, surgical planning, and functional assessment in AF and related disorders. However, most studies remain within early IDEAL stages (1–2b), highlighting the need for multicenter validation, cost-effectiveness analyses, and robust ethical frameworks before widespread clinical implementation.
Anal fistula (AF) is a painful anorectal condition characterized by an abnormal connection between the anal canal and the perianal skin, most commonly resulting from cryptoglandular infection [1, 2]. The annual incidence is estimated at approximately 1 to 2 per 10,000 individuals, with higher rates observed in patients with inflammatory bowel disease, particularly Crohn disease [1]. Patients typically present with pain, discharge, recurrent infections, and potential impairment of continence, all of which can substantially reduce quality of life [3]. Complex AF types, including high transsphincteric, suprasphincteric, or extrasphincteric, horseshoe, and Crohn-related fistulas, pose significant therapeutic challenges because of their intricate anatomy and occult extensions. Recurrence rates in these cases may exceed 40% [1, 4]. Accordingly, the primary treatment objectives are to control sepsis, achieve durable fistula closure, and preserve sphincter function [1].
Anal fistulas rarely resolve spontaneously and almost always require surgical intervention for definitive management. In complex cases, secondary extensions may traverse or extend beyond sphincter planes, making accurate preoperative anatomical mapping essential for selecting an appropriate sphincter-preserving technique. Consequently, several classification systems—most notably the Parks classification and its subsequent refinements—have been developed to characterize tract location, sphincter involvement, and disease complexity [2]. While simple or low fistulas can often be treated effectively with fistulotomy, complex fistulas typically require sphincter-preserving procedures to minimize the risk of postoperative fecal incontinence [5].
Clinical examination and probing have inherent limitations in accurately delineating fistula anatomy, making imaging modalities such as endoanal ultrasonography (EAUS) and magnetic resonance imaging (MRI) essential components of preoperative assessment. EAUS enables real-time evaluation of sphincter integrity, with reported agreement rates of 73% to 100% compared with intraoperative findings [1]. MRI remains the preoperative gold standard, offering superior visualization of fistula tracts and occult extensions compared with clinical examination and ultrasonography; most studies report an accuracy of approximately 90%, although false-positive findings may occur [5]. Following urgent abscess drainage, recurrence rates approach 48% within 1 year, and approximately 40% of patients subsequently develop fistulas over longer follow-up periods [1, 6]. These findings underscore the need for improved diagnostic accuracy and predictive tools.
Advances in computational power and the availability of large datasets have accelerated the development of artificial intelligence (AI), including machine learning and deep learning, across endoscopy and anorectal imaging [7, 8]. Convolutional neural networks (CNNs) have demonstrated utility in enhancing polyp detection, lesion classification, and the reduction of blind spots during endoscopic evaluation [9]. In benign anorectal and pelvic floor disorders, AI has been applied to abscess and fistula imaging, anorectal manometry interpretation, high-definition anoscopy, and outcome prediction [10]. In addition, predictive models integrating clinical, laboratory, and imaging data are being developed to estimate recurrence risk, complications, and treatment outcomes [1012].
Despite these advances, the rationale for integrating AI into AF management requires further clarification. Complex AF exhibits substantial heterogeneity, is difficult to visualize, and is associated with high recurrence rates even after expert surgical management. These characteristics make AF a particularly suitable candidate for computer-assisted imaging, pattern recognition, and predictive modeling. AI has the potential to enhance MRI interpretation, improve detection of occult extensions, support intraoperative decision-making, and enable individualized risk stratification for recurrence. Collectively, these capabilities may facilitate sphincter preservation and reduce the need for repeat interventions.
This review synthesizes current evidence on the application of AI in AF across 4 domains—imaging, intraoperative planning, predictive modeling, and broader anorectal applications—and maps these findings to the IDEAL (idea, development, exploration, assessment, and long-term follow-up) framework. The review further highlights key strengths, limitations, and priorities for future research.
This narrative review identified relevant literature through targeted searches of PubMed, Embase, and Scopus, supplemented by reference chaining. Search terms included combinations of “fistula-in-ano,” “artificial intelligence,” “machine learning,” “deep learning,” “predictive modeling,” “magnetic resonance imaging,” “MRI,” and “surgical planning.” Eligible studies were peer-reviewed English-language articles published from 2010 to 2024 that examined AI applications in AF diagnosis, imaging, management, or outcomes. Abstracts without primary data, non-English studies, and purely technical papers without a clear clinical context were excluded. Because direct AI evidence in anal fistula remains limited, studies involving related anorectal disorders and clinically relevant colorectal surgical applications were also included when they provided potential translational insights into anal fistula diagnosis, functional assessment, surgical planning, or outcome prediction.
Evidence was mapped across 4 domains: (1) preoperative diagnostic imaging; (2) intraoperative surgical planning and assistance; (3) postoperative and predictive modeling; and (4) broader anorectal applications, including anorectal manometry and the EndoFLIP (endoluminal functional lumen imaging probe) procedure. Evidence gaps, including limited external validation, small sample sizes, and reproducibility concerns, were also identified.
Clinical impact was stratified according to the IDEAL framework, which progresses from stage 1 (idea/proof-of-concept) to stage 4 (long-term quality assurance). Stage 2 was subdivided into stage 2a (development: early safety and feasibility) and stage 2b (exploration: efficacy, learning curves, and multicenter data), providing a structured pathway from innovation to comparative effectiveness and long-term monitoring [13, 14]. Stage 3 (assessment) involves prospective comparative evaluation, usually through multicenter studies assessing clinical effectiveness, whereas stage 4 encompasses long-term follow-up and quality assurance after implementation in routine clinical practice. Most included studies corresponded to stages 1–2b; none met criteria for stage 3 or 4.
No formal risk-of-bias tool was applied. Instead, the basic appraisal focused on study design, sample size, and external validation. The emphasis was on providing a broad overview and structured evidence mapping rather than a comprehensive systematic synthesis.
Domain 1: Preoperative imaging

MRI-based AI algorithms (stages 2a–2b)

Deep learning approaches have demonstrated promising performance for fistula and abscess detection. For example, in 50 anorectal cases, a multimodal feature-fusion MRI model outperformed a baseline fully convolutional neural network (FCN), achieving higher clinical detection rates, including similarity of 85.37%, accuracy of 80.02%, and recall of 79.38%; it also demonstrated a higher detection rate than routine MRI (84% vs. 64%) [12]. Importantly, these metrics reflect algorithm-level comparisons rather than comparisons with expert radiologist performance and should therefore be interpreted within the context of early-stage model development. CNN-based models, including MobileNetV2 and ResNet50, have shown high discriminative ability in distinguishing perianal fistulizing Crohn disease from cryptoglandular fistulas, with external validation and superior performance compared with radiologists [15]. Existing evidence derives mainly from single-center studies corresponding to IDEAL stage 2a (development), although a limited number of multicenter datasets with external validation have progressed to stage 2b (exploration). Major gaps include the need for standardized radiologist comparisons, external validation, false-positive reduction, and workflow integration. Given MRI’s strengths and limitations in fistula mapping [5], multicenter validation, standardized reporting, and clear accountability frameworks are essential before widespread adoption.

AI-assisted compressed sensing MRI (stage 2a)

AI-assisted compressed sensing (ACS) MRI addresses prolonged acquisition times. In a series of 51 patients, ACS reduced scan time by more than 50% (from 156 to 74 seconds) while maintaining diagnostic accuracy (88.9%) and improving image quality, including signal to noise ratio and contrast to noise ratio [16]. Clinically, shorter scan times may reduce motion artifacts and increase imaging throughput.
Deep learning FCNs applied to computed tomography (CT) imaging have shown improved lesion segmentation and tissue delineation for perianal abscesses, with potential to support more accurate anatomical assessment [17].
Current evidence remains at the feasibility stage (IDEAL stage 2a), primarily consisting of prospective, single-center research with limited benchmarking and no assessment of cost-effectiveness or patient-reported outcomes. Next steps include pairing ACS with diagnostic classifiers, comparing ACS-MRI with standard MRI protocols, and evaluating costs and downstream testing.
Domain 2: Intraoperative surgical assistance

AI-enhanced 3D modeling for planning (stages 1–2a)

Three-dimensional MRI reconstructions can translate complex perianal anatomy into surgeon-friendly anatomical maps. In Crohn disease, 3D models may improve localization of internal openings, guide seton placement, and support targeted drainage, thereby informing both preoperative planning and intraoperative decision-making [18]. However, the evidence remains limited to IDEAL stages 1–2a, with restricted generalizability and no robust data on cost-effectiveness, recurrence, continence, or other clinically relevant endpoints. Comparative trials are needed to determine whether these tools reduce recurrence, postoperative incontinence, and operative time.

Intraoperative “smart navigation” (stages 1–2a)

AI-derived 3D reconstructions may improve anatomical understanding and surgical planning by enhancing visualization of complex fistula tracts and occult extensions that are difficult to identify on standard imaging. However, evidence for real-time intraoperative navigation and associated clinical outcome benefits remains limited [17]. Deep learning–based CT segmentation may improve delineation of abscess anatomy, but its role in guiding surgical decision-making requires further clinical validation [18]. These technologies remain in early development, corresponding mainly to IDEAL stage 2a, with some studies closer to stage 1. Evidence for operating room integration is limited, and few head-to-head comparisons have been performed. Prospective multicenter studies are needed to assess recurrence, continence, reintervention rates, operative time, costs, and learning curves.
Domain 3: Predictive modeling and postoperative outcomes

Recurrence and outcome prediction (stages 2a–2b)

Reported AF recurrence rates range from 3% to 57%, with an average of approximately 19%. This variability appears to be driven more by anatomical complexity, including multiple tracts, horseshoe configurations, and missed internal openings, than by many patient-level factors [19, 20]. In AF, AI models integrating clinical, laboratory, and imaging features could support risk stratification and individualized counseling [11]. Similar frameworks in colorectal disease have already been used to predict responses to neoadjuvant therapy and postoperative complications, supporting the feasibility of this approach [21, 22].
Most available models are retrospective and rely on internal validation, placing them within IDEAL stages 2a–2b. No studies have met criteria for stage 3, which requires prospective comparative validation and extensive benchmarking against standard tools. This evidence gap creates a substantial risk of overfitting and limits clinical interpretability. Progress will require prospective multicenter validation, transparent benchmarking, and linkage to clinically meaningful outcomes, including recurrence, continence, quality of life, and economic endpoints.

Technique selection and continence preservation (stages 2a–2b)

Because continence preservation is a central goal in AF management, predictive analytics may help match anatomical features to appropriate sphincter-sparing strategies. In a retrospective study, new flatus incontinence occurred in 11.4% of patients (4 of 35) undergoing ligation of the intersphincteric fistula tract and 7.0% of patients (3 of 43) undergoing sphincter-preserving fistulectomy [23]. However, these events were associated with concomitant fistulotomy and were not accompanied by significant differences in validated continence scores. These findings highlight the potential value of decision-support systems that integrate imaging and clinical variables [10, 11]. Current evidence remains observational (IDEAL stages 2a–2b) and has made limited progress toward prospective comparative assessment. Rigorous prospective trials are needed to determine whether AI-assisted technique selection reduces incontinence or recurrence compared with expert judgment alone.
Domain 4: Broader applications of AI in anorectal surgery
Fecal incontinence can be assessed using validated scoring systems, ranging from established instruments such as the Wexner and Vaizey scores to newer frameworks such as Garg’s incontinence score [24, 25]. Beyond AF, AI has shown potential across anorectal physiology and pelvic floor disorders.

AI in anorectal manometry (stages 1–2b)

Machine learning techniques, including gradient boosting, random forests, and support vector machines, applied to high-resolution anorectal manometry (ARM) have achieved more than 80% accuracy in differentiating fecal incontinence from obstructed defecation, with gradient boosting models showing the highest performance (area under the curve [AUC], 0.939) [26]. CNNs may support more consistent and reproducible interpretation of ARM data, although their ability to standardize practice across centers remains unvalidated [26]. The proof-of-concept study by Saraiva et al. [26] corresponds to IDEAL stage 1. Unsupervised learning applied to EndoFLIP and high-definition ARM data has identified distinct manometric profiles with high specificity, supporting diagnostic subtyping [27].

AI in non-AF anorectal imaging contexts (stages 1–2b)

In non-AF anorectal imaging contexts, a deep learning FCN model improved segmentation performance on multislice CT images compared with conventional CNNs, while CT itself demonstrated high diagnostic accuracy (96.67%) for perianal abscess detection [17]. Multimodal MRI deep learning algorithms have reported high detection rates for abscesses, tracts, and internal openings [12]. Even without AI, MRI interpreted by experienced clinicians has demonstrated very high accuracy, including 98.6% for tract detection and 97.7% for identifying internal openings [28]. These findings likely reflect expert-level performance under optimal conditions. By contrast, the lower accuracy reported in early AI models, approximately 80%, likely reflects differences in comparator standards and model maturity. Current AI systems are not designed to replace expert interpretation; rather, they are intended to enhance clinical workflows by improving efficiency, reducing reporting time, and supporting less-experienced operators. A consolidated overview of IDEAL stage mapping across domains, with representative studies, is provided in Table 1 [12, 1518, 21, 22, 26, 27]. Key studies, datasets, performance outcomes, and IDEAL stage classifications are summarized in Table 2 [12, 1517, 21, 22, 28]. The distribution of AI applications across domains and corresponding IDEAL stages is summarized in Fig. 1.
Challenges and shared limitations
Translation into clinical practice is constrained by the predominance of single-center, retrospective studies and the scarcity of multicenter trials, which limit standardization and broader applicability [29]. The lack of large, diverse multicenter datasets contributes to variability in model training, increases the risk of overfitting, and reduces external validity. Infrastructure and cost barriers, particularly in resource-limited settings, may further slow adoption; workflow disruption and clinician skepticism also remain persistent challenges [3033]. Ethical and legal concerns, including data privacy, algorithmic bias, transparency, and accountability, remain insufficiently addressed [34]. Accountability is especially important when AI-assisted misdiagnosis could lead to sphincter injury or unnecessary procedures. Addressing these challenges will improve clinical relevance and support responsible implementation. Across most studies, MRI protocols, annotation quality, and model-training pipelines remain insufficiently standardized, making cross-study comparisons difficult and hindering progression to IDEAL stages 3–4. Consequently, most studies remain in stages 1–2b, with none reaching stage 3 or 4. Although existing studies provide feasibility signals, including improved anatomical mapping and earlier prediction, meaningful translational progress will require external validation, comprehensive bias assessment, prospective multicenter studies, cost-effectiveness analyses, and decision-analytic frameworks. These shared limitations underscore the need for greater methodological rigor and clearer pathways for translating innovation into routine clinical practice.
Future directions

Domain 1: Preoperative imaging

Multicenter validation of MRI- and CT-based AI models for fistula and abscess detection remains a key priority, with particular emphasis on standardized reporting and direct comparison with expert radiologists. Practical evaluation of acquisition-enhancing techniques, such as ACS-MRI, should include cost-effectiveness analyses and patient-reported outcomes [16]. Future studies should adopt standardized imaging protocols, include external test datasets, report calibration metrics, and compare performance against expert interpretation to ensure clinical relevance and generalizability.

Domain 2: Intraoperative surgical assistance

Prospective comparative studies are needed to evaluate AI-assisted 3D reconstruction and deep learning–guided surgical planning or intraoperative navigation, with emphasis on clinically relevant outcomes such as recurrence, continence, reintervention rates, and operative efficiency [18]. This domain should remain distinct from diagnostic applications and should focus specifically on intraoperative decision support. Future research should also assess learning curves, operative time, and downstream healthcare costs while establishing governance frameworks for algorithm-driven surgical decision-making.

Domain 3: Predictive modeling and postoperative outcomes

Future research should focus on developing and externally validating AF-specific predictive models for outcomes such as recurrence, healing, and continence [2123]. Integrating these models into shared decision-making and evaluating their health-economic impact will be essential for clinical translation. Such models should undergo prospective multicenter validation with transparent reporting of calibration, discrimination, and decision-curve analyses to establish robustness and clinical utility.

Domain 4: Broader applications of AI in anorectal surgery

Non-imaging applications, including AI-assisted interpretation of anorectal manometry and other functional assessments, represent an emerging area of research. These methods may improve diagnostic consistency and accessibility, especially in settings with limited subspecialty expertise; however, further prospective validation and standardization are required. Defining this domain around physiological and functional assessment may help avoid overlap with imaging-based diagnostic tools.
Platform and integration
Integrated AI platforms that combine imaging, intraoperative guidance, and predictive analytics could support unified decision-support systems. Hybrid approaches linking MRI-based assessment with intraoperative navigation may improve workflow efficiency and consistency of care. Cloud-based deployment and telemedicine interfaces could further expand access to specialized expertise, although data privacy, equity, and cybersecurity remain essential considerations.
Strengths and limitations
This review provides the first targeted synthesis of AI applications in AF across 4 domains—preoperative imaging, intraoperative assistance, predictive modeling, and broader anorectal applications—organized according to the IDEAL framework (stages 1–4). It demonstrates that current efforts mainly fall within stages 1–2b, identifies common barriers, and outlines the validation, comparison, and outcome studies needed to advance toward stage 4. A major contribution of this review is its explicit alignment of emerging AI tools with IDEAL stage requirements, clarifying both what has been accomplished and what evidence remains necessary for progression toward comparative evaluation and long-term implementation. This structured mapping may help clinicians, researchers, and innovators understand where each modality currently stands on the translational continuum.
By harmonizing data across tables, narrative sections, and IDEAL stage mapping, this review offers a clear framework for interpreting diverse AI studies and guiding future research priorities to improve diagnostic accuracy, surgical planning, and patient-centered outcomes.
This review maps AI applications in AF management within the IDEAL framework, highlighting developmental progress, feasibility signals, and translational gaps across imaging, surgical assistance, predictive modeling, and physiological assessment. A key strength is the consistent categorization across narrative text, tables, and IDEAL stage mapping, which enables clearer interpretation of diverse AI studies. Performance metrics from included studies—including accuracy, similarity indices, and diagnostic parameters—are summarized to facilitate comparative assessment.
However, this review has several limitations. No formal risk-of-bias tool was used, and substantial heterogeneity in study design, patient selection, imaging protocols, and model-training pipelines may have introduced bias and limited generalizability. Most included studies reported only internal validation, such as cross-validation or bootstrapping, whereas external validation in independent or multicenter cohorts was largely absent—an issue specifically highlighted by reviewers.
Although IDEAL staging is useful for structuring innovation pathways, it contains some inherent subjectivity, particularly when distinguishing advanced stage 2b exploration from early stage 3 assessment.
Finally, the lack of comparative outcome data, including recurrence, continence, operative efficiency, quality of life, and cost-effectiveness, limits the ability to determine clinical impact. These limitations reflect the immaturity of the current evidence base and emphasize the need for robust multicenter validation, prospective comparative studies, and long-term surveillance to advance toward IDEAL stage 4.
Conclusions
AI has promising applications in imaging, surgical planning, predictive modeling, and functional assessment for AF. Mapping the current evidence to the IDEAL framework indicates that most studies remain in early stages (stages 1–2b), with none reaching structured assessment (stage 3) or long-term evaluation (stage 4). Although feasibility findings are encouraging, the current evidence is insufficient to support routine clinical use. Further progress will require multicenter validation, standardized methods, comparative trials, and incorporation of patient-centered outcomes, including continence, recurrence, and quality of life. Cost-effectiveness and workflow-integration studies are also essential for real-world adoption. With rigorous evaluation and ethical oversight, AI has the potential to become a reliable tool for improving diagnostic accuracy, supporting sphincter-sparing strategies, and enhancing outcomes in AF management.

Conflict of interest

No potential conflict of interest relevant to this article was reported.

Funding

None.

Author contributions

Conceptualization: all authors; Investigation: all authors; Methodology: all authors; Project administration: VDY, PG; Supervision: all authors; Validation: all authors; Visualization: all authors; Writing–original draft: all authors; Writing–review & editing: all authors. All authors read and approved the final manuscript.

Fig. 1.
Artificial intelligence (AI) in anal fistula: domains and IDEAL (idea, development, exploration, assessment, and long-term follow-up) stage mapping.
ac-2025-01494-0213f1.jpg
Table 1.
IDEAL stage classification of the included studies
Study Study description IDEAL stage Justification
Yang et al. [12] (2021) Deep learning–based MRI features for diagnosing perianal abscess and fistula 2a Diagnostic model development
Zhang et al. [15] (2024) MRI-based deep learning classifier for fistulizing Crohn disease in a multicenter cohort 2b Multicenter validation and exploration
Tang et al. [16] (2023) ACS-MRI in anal fistula 2a Feasibility and early clinical application
Han et al. [17] (2021) Deep learning CT-based detection of perianal abscess 2a Diagnostic model development
Mazaki et al. [21] (2021) AI-based predictive model for anastomotic leakage 2a Retrospective model development with internal validation using 5-fold cross-validation; AUC, 0.766
Bektaş et al. [22] (2022) Systematic review of machine learning in colorectal surgery NA Review article spanning multiple stages
Saraiva et al. [26] (2023) AI in anorectal manometry 1 Initial proof-of-concept demonstration
Zifan et al. [27] (2018) EndoFLIP vs. manometry comparison in fecal incontinence 2b Comparative physiological evaluation
Jeri-McFarlane et al. [18] (2023) 3D modeling for preoperative planning in complex fistula 2a Early prospective feasibility study (n=4); proof of clinical applicability without comparative validation

IDEAL framework: stage 1, idea or proof-of-concept; stage 2a, development; stage 2b, exploration; stage 3, assessment; and stage 4, long-term surveillance.

IDEAL, idea, development, exploration, assessment, and long-term follow-up; MRI, magnetic resonance imaging; ACS, artificial intelligence–assisted compressed sensing; CT, computed tomography; AI, artificial intelligence; AUC, area under the curve; NA, not assignable; EndoFLIP, endoluminal functional lumen imaging probe.

Table 2.
Selected key studies and reported performance metrics
Study Modality/task Study design Key metrics Notes IDEAL stage
Yang et al. [12] (2021) MRI deep learning feature-fusion for abscess/fistula detection Prospective study (50 anorectal cases) Similarity, 85.37% Algorithm vs. FCN comparison; limited clinical validation 2a
Accuracy, 80.02%
Recall, 79.38%
Zhang et al. [15] (2024) MRI deep learning classifier: Crohn disease vs. cryptoglandular fistula Multicenter cohort Internal validation: AUC, 0.962–0.963 External validation; classifier task 2b
External validation: AUC, 0.874–0.885
Superior performance compared with radiologists
Tang et al. [16] (2023) ACS-MRI Prospective comparative study Scan time reduced by >50% (74 sec vs. 156 sec) Acquisition acceleration; feasibility 2a
Increased SNR and CNR
Equivalent diagnostic accuracy (88.9% vs. 88.9%)
Han et al. [17] (2021) CT deep learning FCN for perianal abscess tissue segmentation Single-center comparative study (60 patients and 60 controls) Improved segmentation metrics compared with CNN (Jaccard index, 0.8525 vs. 0.7326; Dice coefficient, 0.8434 vs. 0.7264) Segmentation-focused study; diagnostic accuracy relates to CT, not the AI model 2a
Mazaki et al. [21] (2021) Predictive model for anastomotic leakage in colorectal surgery Retrospective cohort Auto-AI model: AUC, 0.766 Retrospective model development with internal validation using 5-fold cross-validation 2a
Bektaş et al. [22] (2022) Machine learning for surgical outcomes Systematic review Metrics varied by model and outcome Highlights validation gaps NA
Yagnik et al. [28] (2024) Baseline benchmark: MRI accuracy without AI Review Tract detection, 98.6% Clinical benchmark for AI comparison Reference
Internal opening identification, 97.7%

IDEAL framework: stage 1, idea or proof-of-concept; stage 2a, development; stage 2b, exploration; stage 3, assessment; and stage 4, long-term surveillance.

IDEAL, idea, development, exploration, assessment, and long-term follow-up; MRI, magnetic resonance imaging; FCN, fully convolutional neural network; AUC, area under the curve; ACS, artificial intelligence–assisted compressed sensing; SNR, signal to noise ratio; CNR, contrast to noise ratio; CT, computed tomography; CNN, convolutional neural network; AI, artificial intelligence; NA, not assignable.

  • 1. Vogel JD, Johnson EK, Morris AM, Paquette IM, Saclarides TJ, Feingold DL, et al. Clinical practice guideline for the management of anorectal abscess, fistula-in-ano, and rectovaginal fistula. Dis Colon Rectum 2016;59:1117–33. ArticlePubMed
  • 2. Parks AG, Gordon PH, Hardcastle JD. A classification of fistula-in-ano. Br J Surg 1976;63:1–12. ArticlePubMedPDF
  • 3. Abcarian H. Anorectal infection: abscess-fistula. Clin Colon Rectal Surg 2011;24:14–21. ArticlePubMedPMC
  • 4. Hall JF, Bordeianou L, Hyman N, Read T, Bartus C, Schoetz D, et al. Outcomes after operations for anal fistula: results of a prospective, multicenter, regional study. Dis Colon Rectum 2014;57:1304–8. ArticlePubMed
  • 5. Halligan S. Magnetic resonance imaging of fistula-in-ano. Magn Reson Imaging Clin N Am 2020;28:141–51. ArticlePubMed
  • 6. Chaveli Díaz C, Esquiroz Lizaur I, Eguaras Córdoba I, González Álvarez G, Calvo Benito A, Oteiza Martínez F, et al. Recurrence and incidence of fistula after urgent drainage of an anal abscess. Long-term results. Cir Esp (Engl Ed) 2022;100:25–32. ArticlePubMed
  • 7. Min JK, Kwak MS, Cha JM. Overview of deep learning in gastrointestinal endoscopy. Gut Liver 2019;13:388–93. ArticlePubMedPMC
  • 8. Nie Z, Xu M, Wang Z, Lu X, Song W. A review of application of deep learning in endoscopic image processing. J Imaging 2024;10:275.ArticlePubMedPMC
  • 9. Kim NH, Jung YS, Jeong WS, Yang HJ, Park SK, Choi K, et al. Miss rate of colorectal neoplastic polyps and risk factors for missed polyps in consecutive colonoscopies. Intest Res 2017;15:411–8. ArticlePubMedPMCPDF
  • 10. Aleissa M, Osumah T, Drelichman E, Mittal V, Bhullar J. Current status and role of artificial intelligence in anorectal diseases and pelvic floor disorders. JSLS 2024;28:e2024.00007. ArticlePubMedPMC
  • 11. Demirli Atıcı S. Can artificial intelligence be as effective in the treatment of anal fistula as in colorectal surgery? Turk J Colorectal Dis 2022;32:258–9. Article
  • 12. Yang J, Han S, Xu J. Deep learning-based magnetic resonance imaging features in diagnosis of perianal abscess and fistula formation. Contrast Media Mol Imaging 2021;2021:9066128.ArticlePubMedPMCPDF
  • 13. McCulloch P, Altman DG, Campbell WB, Flum DR, Glasziou P, Marshall JC, et al. No surgical innovation without evaluation: the IDEAL recommendations. Lancet 2009;374:1105–12. ArticlePubMed
  • 14. Ergina PL, Barkun JS, McCulloch P, Cook JA, Altman DG; IDEAL Group. IDEAL framework for surgical innovation 2: observational studies in the exploration and assessment stages. BMJ 2013;346:f3011.ArticlePubMedPMC
  • 15. Zhang H, Li W, Chen T, Deng K, Yang B, Luo J, et al. Development and validation of the MRI-based deep learning classifier for distinguishing perianal fistulizing Crohn's disease from cryptoglandular fistula: a multicenter cohort study. EClinicalMedicine 2024;78:102940.ArticlePubMedPMC
  • 16. Tang H, Peng C, Zhao Y, Hu C, Dai Y, Lin C, et al. An applicability study of rapid artificial intelligence-assisted compressed sensing (ACS) in anal fistula magnetic resonance imaging. Heliyon 2023;10:e22817. ArticlePubMedPMC
  • 17. Han S, Yang J, Xu J. Deep learning-based computed tomography image features in the detection and diagnosis of perianal abscess tissue. J Healthc Eng 2021;2021:3706265.ArticlePubMedPMCPDF
  • 18. Jeri-McFarlane S, García-Granero Á, Ochogavía-Seguí A, Pellino G, Oseira-Reigosa A, Gil-Catalan A, et al. Three-dimensional modelling as a novel interactive tool for preoperative planning for complex perianal fistulas in Crohn's disease. Colorectal Dis 2023;25:1279–84. ArticlePubMed
  • 19. Mei Z, Wang Q, Zhang Y, Liu P, Ge M, Du P, et al. Risk factors for recurrence after anal fistula surgery: a meta-analysis. Int J Surg 2019;69:153–64. ArticlePubMed
  • 20. Usta MA. Analysis of the factors affecting recurrence and postoperative incontinence after surgical treatment of anal fistula: a retrospective cohort study. Turk J Colorectal Dis 2020;30:275–84. Article
  • 21. Mazaki J, Katsumata K, Ohno Y, Udo R, Tago T, Kasahara K, et al. A novel predictive model for anastomotic leakage in colorectal cancer using auto-artificial intelligence. Anticancer Res 2021;41:5821–5. ArticlePubMed
  • 22. Bektaş M, Tuynman JB, Costa Pereira J, Burchell GL, van der Peet DL. Machine learning algorithms for predicting surgical outcomes after colorectal surgery: a systematic review. World J Surg 2022;46:3100–10. ArticlePubMedPMCPDF
  • 23. Hong Y, Xu Z, Gao Y, Sun M, Chen Y, Wen K, et al. Sphincter-preserving fistulectomy is an effective minimally invasive technique for complex anal fistulas. Front Surg 2022;9:832397.ArticlePubMedPMC
  • 24. Clemente N, Singhal T, Yagnik VD. Garg incontinence scores: a paradigm shift in assessing fecal incontinence. Glob J Med Pharm Biomed Update 2023;18:19.Article
  • 25. Armstrong DN, Sudoł-Szopińska I, de Parades V, Litta F, Limbert M, James KC. Pankaj Garg: a community doctor to a master innovator to a global icon. Glob J Med Pharm Biomed Update 2023;18:16.Article
  • 26. Saraiva MM, Pouca MV, Ribeiro T, Afonso J, Cardoso H, Sousa P, et al. Artificial intelligence and anorectal manometry: automatic detection and differentiation of anorectal motility patterns: a proof-of-concept study. Clin Transl Gastroenterol 2023;14:e00555. ArticlePubMedPMC
  • 27. Zifan A, Sun C, Gourcerol G, Leroi AM, Mittal RK. EndoFLIP vs high-definition manometry in the assessment of fecal incontinence: a data-driven unsupervised comparison. Neurogastroenterol Motil 2018;30:e13462. ArticlePubMedPMCPDF
  • 28. Yagnik VD, Kumar S, Thakur A, Bhattacharya K, Dawka S, Garg P. Recent advances in the understanding and management of anal fistula from India. Indian J Surg 2024;86:1105–13. ArticlePDF
  • 29. Kenig N, Monton Echeverria J, Muntaner Vives A. Artificial intelligence in surgery: a systematic review of use and validation. J Clin Med 2024;13:7108.ArticlePubMedPMC
  • 30. Pham T. Ethical and legal considerations in healthcare AI: innovation and policy for safe and fair use. R Soc Open Sci 2025;12:241873.ArticlePubMedPMCPDF
  • 31. Hassan M, Kushniruk A, Borycki E. Barriers to and facilitators of artificial intelligence adoption in health care: scoping review. JMIR Hum Factors 2024;11:e48633. ArticlePubMedPMC
  • 32. Maleki Varnosfaderani S, Forouzanfar M. The role of AI in hospitals and clinics: transforming healthcare in the 21st century. Bioengineering (Basel) 2024;11:337.ArticlePubMedPMC
  • 33. Adler-Milstein J, Aggarwal N, Ahmed M, Castner J, Evans B, Gonzalez A, et al. Meeting the moment: addressing barriers and facilitating clinical adoption of artificial intelligence in medical diagnosis. NAM Perspectives; 2022.Article
  • 34. Li YH, Li YL, Wei MY, Li GY. Innovation and challenges of artificial intelligence technology in personalized healthcare. Sci Rep 2024;14:18994.ArticlePubMedPMCPDF

Figure & Data

References

    Citations

    Citations to this article as recorded by  
    • Revisiting the ‘surgeon in the loop’: From deskilling to human–AI collaboration
      Vipul D. Yagnik, Pankaj Garg
      Colorectal Disease.2026;[Epub]     CrossRef
    • Clinical utility beyond detection rates: Interpreting artificial intelligence in FIT‐positive colonoscopy
      Divyesh A. Patel
      Colorectal Disease.2026;[Epub]     CrossRef

    • Cite this Article
      Cite this Article
      export Copy Download
      Close
      Download Citation
      Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

      Format:
      • RIS — For EndNote, ProCite, RefWorks, and most other reference management software
      • BibTeX — For JabRef, BibDesk, and other BibTeX-specific software
      Include:
      • Citation for the content below
      Artificial intelligence in anal fistula: mapping evidence to IDEAL stages
      Ann Coloproctol. 2026;42(3):264-272.   Published online June 25, 2026
      Close
    • XML DownloadXML Download
    Figure
    • 0
    Artificial intelligence in anal fistula: mapping evidence to IDEAL stages
    Image
    Fig. 1. Artificial intelligence (AI) in anal fistula: domains and IDEAL (idea, development, exploration, assessment, and long-term follow-up) stage mapping.
    Artificial intelligence in anal fistula: mapping evidence to IDEAL stages
    Study Study description IDEAL stage Justification
    Yang et al. [12] (2021) Deep learning–based MRI features for diagnosing perianal abscess and fistula 2a Diagnostic model development
    Zhang et al. [15] (2024) MRI-based deep learning classifier for fistulizing Crohn disease in a multicenter cohort 2b Multicenter validation and exploration
    Tang et al. [16] (2023) ACS-MRI in anal fistula 2a Feasibility and early clinical application
    Han et al. [17] (2021) Deep learning CT-based detection of perianal abscess 2a Diagnostic model development
    Mazaki et al. [21] (2021) AI-based predictive model for anastomotic leakage 2a Retrospective model development with internal validation using 5-fold cross-validation; AUC, 0.766
    Bektaş et al. [22] (2022) Systematic review of machine learning in colorectal surgery NA Review article spanning multiple stages
    Saraiva et al. [26] (2023) AI in anorectal manometry 1 Initial proof-of-concept demonstration
    Zifan et al. [27] (2018) EndoFLIP vs. manometry comparison in fecal incontinence 2b Comparative physiological evaluation
    Jeri-McFarlane et al. [18] (2023) 3D modeling for preoperative planning in complex fistula 2a Early prospective feasibility study (n=4); proof of clinical applicability without comparative validation
    Study Modality/task Study design Key metrics Notes IDEAL stage
    Yang et al. [12] (2021) MRI deep learning feature-fusion for abscess/fistula detection Prospective study (50 anorectal cases) Similarity, 85.37% Algorithm vs. FCN comparison; limited clinical validation 2a
    Accuracy, 80.02%
    Recall, 79.38%
    Zhang et al. [15] (2024) MRI deep learning classifier: Crohn disease vs. cryptoglandular fistula Multicenter cohort Internal validation: AUC, 0.962–0.963 External validation; classifier task 2b
    External validation: AUC, 0.874–0.885
    Superior performance compared with radiologists
    Tang et al. [16] (2023) ACS-MRI Prospective comparative study Scan time reduced by >50% (74 sec vs. 156 sec) Acquisition acceleration; feasibility 2a
    Increased SNR and CNR
    Equivalent diagnostic accuracy (88.9% vs. 88.9%)
    Han et al. [17] (2021) CT deep learning FCN for perianal abscess tissue segmentation Single-center comparative study (60 patients and 60 controls) Improved segmentation metrics compared with CNN (Jaccard index, 0.8525 vs. 0.7326; Dice coefficient, 0.8434 vs. 0.7264) Segmentation-focused study; diagnostic accuracy relates to CT, not the AI model 2a
    Mazaki et al. [21] (2021) Predictive model for anastomotic leakage in colorectal surgery Retrospective cohort Auto-AI model: AUC, 0.766 Retrospective model development with internal validation using 5-fold cross-validation 2a
    Bektaş et al. [22] (2022) Machine learning for surgical outcomes Systematic review Metrics varied by model and outcome Highlights validation gaps NA
    Yagnik et al. [28] (2024) Baseline benchmark: MRI accuracy without AI Review Tract detection, 98.6% Clinical benchmark for AI comparison Reference
    Internal opening identification, 97.7%
    Table 1. IDEAL stage classification of the included studies

    IDEAL framework: stage 1, idea or proof-of-concept; stage 2a, development; stage 2b, exploration; stage 3, assessment; and stage 4, long-term surveillance.

    IDEAL, idea, development, exploration, assessment, and long-term follow-up; MRI, magnetic resonance imaging; ACS, artificial intelligence–assisted compressed sensing; CT, computed tomography; AI, artificial intelligence; AUC, area under the curve; NA, not assignable; EndoFLIP, endoluminal functional lumen imaging probe.

    Table 2. Selected key studies and reported performance metrics

    IDEAL framework: stage 1, idea or proof-of-concept; stage 2a, development; stage 2b, exploration; stage 3, assessment; and stage 4, long-term surveillance.

    IDEAL, idea, development, exploration, assessment, and long-term follow-up; MRI, magnetic resonance imaging; FCN, fully convolutional neural network; AUC, area under the curve; ACS, artificial intelligence–assisted compressed sensing; SNR, signal to noise ratio; CNR, contrast to noise ratio; CT, computed tomography; CNN, convolutional neural network; AI, artificial intelligence; NA, not assignable.


    Ann Coloproctol : Annals of Coloproctology Twitter Facebook
    TOP