Cleft Lip & Palate Research

A synthesis of 33 papers, seven open research questions, and a practical guide to doing work that gets published.

See the open questions

This page is a working map of the cleft lip and palate literature — what is settled, what is genuinely unresolved, and where a new researcher can contribute something real rather than another underpowered case series.

It comes from reading 33 papers spanning 2012–2026: full text for the surgical and presurgical-orthopedics cluster, abstracts plus discussion and limitations sections for the rest. Every number quoted below is traceable to a specific paper, listed in the reference table at the end.

1. The state of the field

The most important thing to notice about this literature is not any single result. It is that the strongest papers in the field are mostly negative or inconclusive, and they fail in the same way.

Four serious, well-conducted reviews, all arriving at the same place: the primary literature is too heterogeneous, too small, and too inconsistently reported to answer the questions the field actually cares about.

Why this is good news for you

A field whose reviews all conclude "the evidence is too heterogeneous" is a field with structural problems that are solvable by careful work — standardizing a definition, testing a confounder, re-analyzing existing data with better methods. Those projects do not require a laboratory, a patient panel, or a grant. They require reading closely and thinking clearly, which is exactly what you can do right now.

2. What the literature covers

The 33 papers fall into nine clusters. The mass is heavily concentrated in presurgical orthopedics and surgical protocol comparison.

ClusterPapers
Presurgical infant orthopedicsPapadopoulos 2012, Hosseini 2016, Górska 2022, Di Blasio 2024, Rabal-Soláns 2024, Bühling 2025, Machado 2026
Surgical protocol & growthvan Roey 2025, Kleijnen 2023, Wu 2023, Demiröz 2021, Ameer 2023
Speech & velopharyngeal functionAlighieri 2025, Tahmasebifard 2025, Liu 2026, Du 2022, Tosun 2024
Feeding & growthKinter 2026, Ormanidou 2026, Williams 2025, Kotlarek 2026
Dental anomaliesPradhan 2020, Batwa 2018, Hamid 2025
Patient-reported outcomesMichael 2022, Kabuyaya 2024, Ruiz-Guillén 2021
Epidemiology, etiology, cultureDuan 2023, Eniyew 2026, Hasanuddin 2025
Perioperative careSuleiman 2024, Grabar 2023
AI & digital methodsÖztürk 2026

3. Seven open research directions

Each of these comes from a specific gap that the authors themselves named and then did not fill. That matters: a question an established author has publicly flagged as unanswered is a question you can pursue without having to argue that it is worth asking.

Direction 1

Stop asking whether presurgical orthopedics works — ask for whom

The tension. Four meta-analyses across fourteen years (Papadopoulos 2012, Hosseini 2016, Di Blasio 2024, Rabal-Soláns 2024) find no significant effect of presurgical orthopedics on arch dimensions, feeding, or growth. But Bühling (2025), the only prospective interventional study in this set, finds active plates outperform passive ones — 5.85 mm versus 3.57 mm of cleft width reduction — and names its own key limitation:

"the lack of numerical or categorical guidelines for the allocation of patients into the active or passive plate groups."

Bühling also reports that baseline cleft size did not predict how much change occurred. That is either a real and surprising finding or an artifact of a 20-patient sample, and it matters enormously which.

The study. Derive and externally validate a morphology-based rule for allocating infants to active, passive, or no plate. Use archived neonatal casts — every cleft center has decades of them in cupboards. Test the treatment-by-morphology interaction rather than the main effect.

Why it is feasible. Framed as heterogeneous treatment effect estimation, this sidesteps the ethical objection Machado raises about withholding treatment from newborns: you are comparing two accepted treatments, not treatment against nothing.

Direction 2

Test whether the protocol literature is secretly measuring access to orthodontic care

The tension. van Roey's most interesting paragraph is buried in his discussion. Dental arch relationships were worse under one-stage palatoplasty — but 67% of those patients came from Asia and South America, where cleft care is often not covered by insurance, against 60% of Oslo-protocol patients from Europe and Oceania. He then cites Southall: orthodontic treatment alone shifts the GOSLON score by 1.1 points, which is larger than the between-protocol differences the field has been arguing about for decades.

He raises the hypothesis. He does not test it.

The study. Meta-regression across his own 156 study groups, with country income, insurance coverage, and reported orthodontic exposure as moderators of arch relationship outcomes.

Why it is feasible. The data are already extracted and published in his supplementary files. This is weeks of work, not years, and no patients are involved. The conclusion would be outsized: it would tell the field how much of its central controversy is an artifact of who can afford a retainer.

Direction 3

Build the preoperative VPI prediction model that Liu (2026) tried and failed to build

The tension. Liu compared pharyngeal anatomy in isolated cleft palate against Pierre Robin sequence and concluded there was no difference in need ratio. But the cleft palate group had a median age of 23.7 months and the Pierre Robin group 3.8 months. The paper's own discussion concedes the age gap "likely explains" the differences found. As a comparison it is uninterpretable.

Meanwhile Tahmasebifard (2025) has the right method — MRI, 105 individuals with velopharyngeal insufficiency against 45 controls, closure height 3.7 mm lower in the VPI group — but is cross-sectional, underpowered for subgroups, and explicitly flags that "the timing and type of palatal surgery might also influence the height of VP closure. Future studies should evaluate the impact of these factors."

The study. A prospective, age-standardized preoperative imaging cohort. Measure need ratio, closure height, velar length, and cleft width at a fixed preoperative timepoint; follow to speech assessment at five years. Output: a calibrated risk score.

Why it matters. It would let surgeons match the palatoplasty technique to the anatomy rather than to habit, and it converts the Pierre-Robin-versus-isolated-cleft debate into a covariate.

Direction 4

Close the nutrition → surgical timing → complication loop

The tension. Four papers touch this chain and none connect it.

  • Kinter (2026) — children with clefts sit near −1 SD weight-for-age; 10–50% have failure to thrive; data past 12 months is sparse. Cites a finding that length predicts fistula better than weight, then says "further studies are required to identify the metrics with the greatest clinical relevance."
  • Williams (2025) — Pierre Robin sequence, poor growth, and Hispanic ethnicity independently predict delayed palate repair across 17 teams.
  • Wu (2023) — in 980 palatoplasties, poor healing tracks cleft width, operative time, and surgeon seniority. Nutritional status was not among the recorded variables.
  • Ormanidou (2026) — lower birthweight in infants with clefts, stratified by country income.

The study. Does anthropometric status at the time of palatoplasty predict fistula and velopharyngeal insufficiency, independent of cleft width and surgeon experience? And the harder question: is delaying surgery to improve growth actually protective, or does it simply trade a healing benefit for a speech cost? Wu's dataset is nearly the right shape; it needs weight and length added.

The spin-off. The ethnicity finding deserves its own paper. Williams says only that "future research should investigate social determinants of health that may predict timing of palate repair." Nobody has.

Direction 5

Test whether CLEFT-Q means the same thing in every country using it

The tension. This is the most methodologically distinctive question in the set, and it has teeth because ICHOM has adopted CLEFT-Q as a global standard.

  • Michael (2022), in Nigeria — after palatoplasty her patient exceeded normative CLEFT-Q values for speech function, speech distress, and psychological function, but school and social function stayed below the 95% confidence interval. Her conclusion: "the cleft Q may need further validation in our setting."
  • Kabuyaya (2024) then uses the French CLEFT-Q in the DRC without addressing this, and concedes the study "did not take into account other potential factors… such as socioeconomic conditions."
  • Ruiz-Guillén (2021) finds gender moderates improvement specifically in the social function domain.
  • Hasanuddin (2025) documents, across 13 countries, that clefts are attributed to curses, eclipses, and ancestral spirits, with consequences running to stigmatization and infanticide.

The study. Formal measurement invariance and differential item functioning testing of the CLEFT-Q psychosocial subscales, with Hasanuddin's cultural-belief taxonomy as the grouping variable.

Why it matters. If the instrument is not invariant, every cross-country CLEFT-Q comparison — including the normative values everyone benchmarks against — rests on shaky ground. That is an important result whichever way it comes out.

Direction 6

Replace "cleft width" with automated 3D shape phenotyping

The tension. Du (2022) ends by asking for exactly this:

"more studies are required where data mining and machine learning techniques can be implemented to quantitatively determine the characteristics of different subtypes of the malformation… the computational analysis of the 3D model where topological features such as area and length of the palatal bony structure can be automatically extracted."

Bühling (2025) has already shown that digital measurement on 3D scans agrees with manual measurement on plaster casts, so the method is validated.

Almost every quantitative claim in this literature compresses a complex three-dimensional deformity into one or two numbers — cleft width, cleft palate index, alveolar cleft width. That is very likely a major source of the heterogeneity that Rabal-Soláns and Machado both blame for their inconclusive results.

The study. Statistical shape modelling of neonatal cleft maxillae from digitized archival casts, then use the shape coordinates rather than the scalars as predictors in Directions 1 and 4.

Why it is feasible. This is the enabling method for two other directions, and the input data is already sitting in storage at every major center.

Direction 7

Two quicker, cleaner wins

7a. A reporting-bias-corrected estimate of VPI and fistula rates. van Roey's funnel plot shows that studies reporting high rates of velopharyngeal insufficiency are underrepresented, and he states plainly that "in practice, the incidences of VPI are likely to be higher." He also notes that definitions of oronasal fistula are inconsistent or entirely absent across studies. So: apply selection models to his extracted data to estimate the true rates, and run a Delphi process to fix a definition. He even proposes the long-term mechanism — registries let complication rates be reported without tracing back to individual institutions, "which may reduce reluctance to report." That is an empirically testable claim about how researchers behave.

7b. An enhanced-recovery protocol with healing, not pain, as the primary outcome. Grabar (2023) found 61% of cleft surgeons never use a standardized perioperative protocol, but 68% would adopt one. Suleiman (2024) already supplies the evidence-based analgesia regimen and flags dexamethasone and pre-incisional infiltration as unresolved. Wu (2023) supplies the missing link: agitation and crying drive wound dehiscence and fistula. Nobody has run a cleft enhanced-recovery trial with fistula as the endpoint rather than a pain score — which is the outcome surgeons actually care about, and therefore the reason they would adopt the protocol.

If you pick two, pick these

Direction 2 — the access-as-confounder meta-regression. It is cheap, it uses already-published data, and it could reframe the field's central debate.

Direction 1 — presurgical orthopedics as an effect-modification problem. Seven papers in this collection have circled it for fourteen years without landing on it.

4. A contradiction worth resolving

The literature contains a direct conflict on speech therapy that nobody has addressed.

They may be defining the comparator differently — Kleijnen's is "purely phonological," Alighieri's is combined — but that distinction is carrying a great deal of weight and has never been stated explicitly in print.

Alighieri's trial has 14 participants. Given that van Roey notes children in the Scandcleft trial who attended extensive speech therapy still had poor consonant accuracy, and that Alighieri's own earlier work documents parental "treatment fatigue and dropout," an adequately powered trial stratified by error type is overdue.

A short, publishable first project: a narrative review that reconciles the two, making the definitional difference explicit. Reconciling a contradiction between two respected sources is a real contribution, and it is achievable without collecting any data.

5. How to become exceptional in this field

Most people entering cleft research write a retrospective chart review of their institution's last hundred cases, publish it in a mid-tier journal, and add nothing. The reviews above explain exactly why that fails: the field is drowning in small, heterogeneous, inconsistently reported single-center studies. Adding one more makes the problem marginally worse.

The way to stand out is to work on the structural problems instead. Here is how.

Read limitations sections first

This is the single highest-leverage habit in research reading. Abstracts tell you what authors want to claim; limitations sections tell you what they actually know. Every one of the seven directions above came from a limitations or discussion paragraph where an author named a gap and walked away from it.

For each paper, keep three columns: known, uncertain, unknown. The third column is your research pipeline.

Play to your actual position

You do not currently have a patient panel, a laboratory, or a grant. That rules out some research — and it points directly at the kind that is both achievable and unusually valuable right now: secondary-data work. Meta-regression, reanalysis of published data, psychometric validation, systematic review, methodological critique.

Directions 2, 5, and 7a require nothing but published papers, a laptop, and statistical care. They are also among the most consequential ideas on this page. That is not a coincidence: because they require no clinical access, nobody with clinical access has bothered to do them.

Learn the four tools the good papers all use

Every strong paper in this collection uses the same small toolkit. Learn it once and you can read — and write — at the top of the field.

PRISMA reporting standard for systematic reviews GRADE rating certainty of evidence ROBINS-I / RoB 2 risk-of-bias assessment R: meta, metafor meta-analysis and meta-regression

van Roey's paper is the best worked example: read its methods section closely and you will see all four applied properly. Then register your own protocol on PROSPERO before you start — it is free, it takes an afternoon, and it signals seriousness to any mentor you approach.

Approach mentors with a protocol, not a request

"Can I help with your research?" gets a polite no, because it creates work for the recipient. "I have drafted a protocol to test whether the association between surgical protocol and arch relationship is confounded by orthodontic access, using the 156 study groups in your 2025 meta-analysis — would you be willing to look at it?" is a different message entirely. It offers labour rather than asking for it, and it demonstrates that you read the paper properly.

The authors in this collection who are most worth writing to are the ones who publicly named gaps they had no plans to fill: van Roey, Bühling, Du, Kinter, Williams. Reference their exact sentence.

Get into the field's infrastructure early

Cleft care is unusually well organized internationally, and the organizations are where the data and the collaborators are.

  • ACPA (American Cleft Palate–Craniofacial Association) — the professional home of the field in North America, and publisher of The Cleft Palate–Craniofacial Journal. Its annual meeting takes abstracts; a well-executed secondary analysis is very much within scope for a poster.
  • ERN CRANIO — the European reference network behind van Roey's work, and the registry he proposes as the fix for under-reported complications.
  • CORNET — the Cleft Outcomes Research Network consortium behind the Williams multisite study.
  • ICHOM — publishes the standard outcome set for cleft care, which is the reference point for any outcomes work.

Being a named, familiar participant in these communities compounds over years. Start now, while you have reading time.

Pick one question and go deep

The temptation is to work across surgery, speech, feeding, and psychosocial outcomes at once. Resist it. The researchers who matter in this field are known for one thing — van Roey for protocol comparison, Kinter and McKinney for growth, Alighieri for speech intervention, Perry for velopharyngeal imaging. Depth is what generates invitations, collaborations, and eventually your own trial.

Given your goal of operating on children with clefts, the natural centre of gravity is the surgical cluster: Directions 1, 2, 3, and 7b all sit there, and all of them connect to decisions you will personally be making in an operating room in a few years.

Write continuously, not eventually

Keep a running document of every gap you find, with the exact quotation and citation. Within six months of disciplined reading you will have something most junior researchers never assemble: a defensible, evidence-based map of what the field does not know. That document is the raw material for a review article, a research statement, an interview answer, and eventually a grant.

A realistic first year

  • Months 1–3. Read the surgical and presurgical-orthopedics cluster in full, limitations first. Build the known/uncertain/unknown table. Learn PRISMA and GRADE by reading van Roey's methods closely.
  • Months 3–6. Write the speech-therapy reconciliation review from section 4 — short, achievable, genuinely useful. Register a PROSPERO protocol for Direction 2 or 7a.
  • Months 6–12. Execute the secondary analysis. Submit an abstract to the ACPA annual meeting. Send the protocol to two of the authors who named the gap.

That path produces a first-author publication, a conference presentation, and two senior collaborators — without requiring a lab, a grant, or a single patient.

6. Key figures and sources

Every claim on this page traces to one of these papers.

PaperFigures cited above
van Roey 2025162 articles / 156 study groups; 4,040 one-stage + 1,632 Oslo + 791 delayed-closure patients. VPI: Oslo 24%, one-stage 14%, delayed 9%. Fistula: Oslo 7%, one-stage 10%, delayed 20%. Arch-relationship sample 32.5% Asia and 34.3% South America (one-stage) against 59.8% Europe/Oceania (Oslo). Southall: 1.1-point GOSLON gain from orthodontics.
Kleijnen 202329 clinical questions; systematic-review evidence found for 8; 21 unanswered.
Rabal-Soláns 20245 studies. Alveolar cleft width mean difference −3.06 (95% CI −8.03 to 2.70), I² = 99%. Posterior cleft width −0.88 (−2.06 to 0.30), I² = 89%.
Machado 2026207 studies mapped; nasoalveolar molding most frequent; predominance of case reports and case series.
Bühling 202520 unilateral cleft lip and palate + 11 isolated cleft palate controls. Mean cleft width reduction 5.05 mm; active 5.85 vs passive 3.57 mm (p = 0.024). Baseline cleft size did not affect change. Digital and manual measurement agreed.
Wu 2023980 palatoplasties. 84.2% normal healing, 11.4% delayed, 4.4% fistula. Risk factors: surgeon seniority, widest cleft width, cleft palate index, operation time.
Kinter 2026143 eligible studies, 46 with comparable estimates. Mean weight-for-age z-score near −1 SD; 10–50% failure to thrive or underweight. Sparse data past 12 months.
Williams 2025414 infants (268 cleft palate, 146 with Pierre Robin sequence) across 17 teams. Repair at 13.55 vs 12.05 months. Delay predictors: Pierre Robin sequence, Hispanic ethnicity, poor growth.
Tahmasebifard 2025MRI; 105 with VPI vs 45 controls. Closure height 3.7 mm lower in VPI. No sex effect; underpowered for subgroups.
Liu 202629 cleft palate vs 18 Pierre Robin sequence. Need ratio 0.728 vs 0.748 (not significant), both above the normal 0.6–0.7. Median age 23.7 vs 3.8 months — the confound.
Du 202234 submucous cleft palate with cleft lip. No correlation between cleft lip severity and palatal bony defect. 94% male. Calls for machine learning on 3D models.
Alighieri 202514 children (7 per arm). One-week intensive, two sessions daily. Combined phonetic–phonological outperformed motor–phonetic; maintained at 3 months only in the combined arm.
Grabar 202331 of 102 surgeons responded (30.4%). 61.3% never use a standardized protocol; 67.7% willing to adopt one.
Suleiman 202419 randomized trials + 4 systematic reviews. Recommends suprazygomatic maxillary nerve block plus dexmedetomidine and acetaminophen/NSAIDs. Dexamethasone and pre-incisional infiltration unresolved.
Michael 2022Nigeria; adult with Veau II cleft. Post-operative scores exceeded normative values for speech and psychological function; school and social function remained below the 95% CI.
Kabuyaya 2024DRC; 43 patients aged 8–29; French CLEFT-Q before and 3 months after surgery.
Hasanuddin 202514 articles, 13 countries, 4 continents. Natural versus supernatural attribution of cleft causation.
Ormanidou 20265,651 articles screened, 296 reviewed in full. Birthweight and low birthweight by cleft type, comorbidity, and country income.
Kotlarek 202656.6% of academic programs include no cleft-feeding content; 81.6% of providers trained on the job.
Öztürk 202630 mothers; single-group pre/post with immediate post-test. Largest knowledge gain in complication monitoring, which had the lowest baseline.
Pradhan 2020208 panoramic radiographs. 90.4% had at least one anomaly; agenesis 77.9%; lateral incisor most affected (65.2%).
Duan 20231,051 Tibetan patients. Cleft lip 36.44%, cleft palate 13.32%, both 50.24%. Sex ratios differ by phenotype; syndromic 3.43%.
Eniyew 2026544 children; Bayesian bivariate multinomial model. Maternal nutrition, folate, smoking, rural residence.
Batwa 201896 unilateral complete cleft records. Missing teeth affected dental but not skeletal characteristics; overjet reduced with multiple missing teeth.

Get in touch

I am very interested in research collaborations and initiatives to help children born with cleft lip and palate.

Email me