Health NewsNewsTech News

A Light Box, Not a Diagnosis: Iraqi and Australian Engineers Matched Tongue Color to Illness in 58 of 60 Hospital Cases

The number that traveled around the world was not the number in the hospital. Engineers at Middle Technical University in Baghdad and the University of South Australia trained machine-learning models on 5,260 tongue photographs and reported that the best color-naming systems cleared 98 percent accuracy on that training task. When the same idea was pointed at 60 real patients in Iraqi teaching hospitals, the match to a charted illness was 58 out of 60, or about 96.6 percent. Those are different claims. One says a model can name a color under conditions the researchers controlled. The other says that name lined up with a diagnosis already in a medical record. Press offices, including UniSA’s August 13, 2024 release, folded both into a single promise: say “aah” and get a diagnosis on the spot. The paper itself, “Tongue Disease Prediction Based on Machine Learning Algorithms,” published in the journal Technologies in July 2024, is narrower than that sentence.

The AEGIS Alliance read the paper, the university release, and the follow-on notes that appeared through 2025. Senior author Ali Al-Naji, an adjunct associate professor tied to both MTU and UniSA, has been plain about the clinical associations the group used. A yellow tongue showed up with diabetes. A purple tongue with a thick, greasy coating showed up with some cancers. Acute stroke patients in the set presented with an unusually shaped red tongue. A white tongue tracked with anemia. Severe COVID-19 tracked with a deep red. Indigo or violet coloring tracked with some vascular problems, gastrointestinal illness, or asthma. Co-author Javaan Chahl, joint chair of sensor systems at UniSA, has added a finer COVID scale in later explanations: faint pink in milder illness, crimson in moderate cases, deep red when the infection was serious. None of those pairings is a blood test. They are patterns in a small bedside sample.

The sample is the part most headlines skipped. Two teaching hospitals, in Dhi Qar and Mosul, supplied the 60 photographs in 2022 and 2023. The images were not scraped from social media. Patients put their heads into an LED-lit box. A camera sat about 20 centimeters from the tongue. Of six algorithms the team tried, Extreme Gradient Boosting was the strongest at naming color. Color naming is a classification problem. Diagnosis is a clinical one. The leap in the study is that the color the model named agreed with the chart in 58 of those 60 controlled photographs. Sixty people is a pilot. It is not a multicenter trial, and it is not a device a regulator has cleared for a clinic.

A researcher demonstrates a camera capturing tongue photographs for machine-learning screening.
A researcher demonstrates a camera capturing tongue photographs for machine-learning screening. (Middle Technical University)

The box of lights is the actual invention

Earlier tongue-imaging papers failed in ordinary rooms. A phone in a kitchen changes white balance every time a window or a lamp moves. A “yellow” tongue becomes a tungsten bulb. Chahl’s kiosk is the unglamorous fix: lock the light, lock the distance, then let the classifier work. Take the box away and the 98 percent figure is a lab result looking for a bathroom mirror. The long-term pitch, which Chahl has repeated since the 2024 release, is a smartphone version people could use at home as a screen, not as a doctor. That pitch is a product roadmap. It is not a result the July 2024 paper delivered.

The work did not stop at that paper. In June 2025 Al-Naji and Chahl described preliminary results from a Streamlit web application that reads tongue shape and color and sets the reading next to both traditional Chinese medicine categories and Western-medicine frames. The group has also run the YOLO object-detection algorithm over roughly 750 internet images to look for ulcers and cracks. Separate 2025 work, including a paper in Scientific Reports, tested tongue photographs as a non-invasive clue in coronary artery disease. Scientific American returned to the original study in October 2025 and quoted Al-Naji on a tighter crop: he was working on reading the center and the tip rather than the whole surface. Each of those projects is a different claim. A browser demo is not a hospital trial. An internet scrape of dramatic tongues is not the LED kiosk. A coronary-artery model is not the seven-color sorter.

Traditional Chinese medicine is the ancestor the press release wants readers to picture. Practitioners have inspected tongues for about two millennia, and the World Health Organization has folded some TCM diagnoses into the International Classification of Diseases. That institutional nod does not settle the scientific argument. Plenty of physicians still treat tongue color as a soft sign at best: useful when it is extreme, useless when a patient wants a number. Al-Naji’s own line, carried by the university, is that color, shape, and thickness “can reveal a litany of health conditions.” “Can reveal” is not “will diagnose.” The AEGIS Alliance has watched other consumer health tools make the same slide, from a clever sensor to a wellness advertisement, on the health and technology desks.

What a photograph is structurally unable to see

The researchers warn that many diseases leave no mark on the tongue at all. A model trained on seven colors — red, yellow, green, blue, gray, white, and pink — will always return one of those colors. It will not find a silent tumor, a lab value, or a history the patient did not offer. It will also inherit the lighting and the population it was shown. The bedside photographs come from two hospitals in Iraq. Skin tone, diet, smoking, antibiotics, and coffee all change a tongue. A classifier that has not been audited across those variables will look brilliant in the room where it was built and ordinary in the next country. Internet images used for the ulcer work add a second bias: people photograph what looks alarming, then a detector learns to find alarm.

Privacy is the problem the engineering paper treats lightly and a newsroom should not. A tongue photograph is a biometric. If the frame is sloppy it includes a face. A home app that uploads that image to a cloud model is a medical data pipe dressed as a curiosity. Someone has to answer who stores the picture, how long it lives, whether it can train the next model, and whether an insurer or an employer can ever demand it. Those questions sit outside the accuracy table. They are the reason a smartphone “screen” is a regulatory product, not a graduate-school demo. Pulse oximeters already guess oxygen from two wavelengths of light. Camera apps already guess heart rate from a fingertip. Tongue color is older than both. What this group added is a classifier plus a lighting protocol. What it has not added is a cleared device, a bias audit, or a trial large enough that a hospital would budget for the kiosk.

There is a useful version of this research, and it is modest. Primary-care waits in many countries are measured in weeks. A cheap, well-lit photograph that says “this is not the pink a chart expects” could pull a high-risk patient into a real exam sooner than a waiting list would. That is triage. It is not a diagnosis delivered by a notification. Readers who want the methods can start with the open Technologies paper and with UniSA’s own description of the work. Readers who want the institutional fights around medical claims can stay with The AEGIS Alliance science coverage, including reporting on how a mouth finding becomes a heart argument long before anyone agrees the link is proved. This article is general information, not medical advice. If a tongue looks wrong, the next step is still a licensed clinician.

The honest split is easy to lose once a headline has 98 percent in it. Inside the light box, the models were very good at naming colors, and in 58 of 60 hospital cases that name agreed with an illness already written down. Outside the box, in a bathroom, on a phone, across a population the training set never saw, the result has not been shown. Al-Naji and Chahl built a careful optical trick and a classifier. They did not build a clinic that fits in a pocket. Until a locked white-balance phone trial, on a much larger and more varied group of patients, is published and reviewed, the figure that belongs in public is 58 out of 60 under LEDs — and the sentence that belongs next to it is that a photograph is a flag, not a physician.

Kyle James Lee
Majority Owner of The AEGIS Alliance. I studied in college for Media Arts, Game Development. Talents include Writer/Article Writer, Graphic Design, Photoshop, Web Design and Development, Video Production, Social Media, and eCommerce.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button
Signup for our news and memes newsletters! 

Newsletter Form

Lists
close-link