By Arkmeds
Artificial intelligence can help clinical engineering teams organize information and identify issues that need attention. For this support to be useful, the team needs to know where the data comes from, what task the tool performs and who reviews the result. Start with a specific problem, such as making service history easier to search or helping complete records. Define how the result will be evaluated.
In its introduction to Mark II, Arkmeds describes AI applications that support data analysis, suggest entries and flag inconsistencies. These functions process information within the software. Measurements are performed by the instruments used during the service. Read about the Mark II integration approach, in Portuguese.
Start with a task that can be reviewed
Describe the task in simple terms. What goes into the system? What should it produce? Who receives the response? What evidence will show whether it was useful?
A clearly defined task makes evaluation easier. For example, a team may want a text suggestion to preserve the information in a work order and help identify gaps for review. In that case, success includes factual accuracy and clarity about what is still missing.
Suggested text should be checked by someone who understands the service and can verify the source information. A fluent response may contain an inappropriate interpretation or a fact that is absent from the record.
Quality starts with the asset record and the work order
Incomplete or ambiguous data makes later analysis more difficult. To prepare the data, start with problems the team can already recognize: duplicate assets, generic descriptions, inconsistent dates, fields with missing units and closed work orders that do not explain what was done.
As an initial review, check whether the records answer these questions:
- Which device was serviced?
- Why was the work needed?
- What was observed and done?
- What results and limitations were recorded?
- What remains outstanding, and who is following it up?
These questions help assess the information available; they do not require every work order to become a lengthy document. A concise, verifiable description can be more useful than several paragraphs without context. Length alone is not a measure of quality.
Distinguish measurements, automation and suggestions
When reviewing a solution, distinguish three sources of information: values obtained by instruments, fields populated through an integration and content suggested by AI. Each source calls for a different type of check.
For a measured value, the team must preserve its connection to the instrument, the asset and the procedure. For an automatically populated field, confirm that the data was placed in the correct field and work order. For an AI suggestion, check whether the content matches the available evidence and whether the interpretation is appropriate.
Identifying the source of incorrect information helps investigate the discrepancy and correct the process that produced it.
Why a convincing reference still needs to be checked
A study published in 2023 examined a specific problem in text generation: fabricated bibliographic references. During the first week of April that year, the researchers asked GPT-3.5 and GPT-4 to produce literature reviews on 42 topics and checked 636 references in 84 texts. In that sample, 55% of GPT-3.5 references and 18% of GPT-4 references did not correspond to real works. The checks sought to distinguish nonexistent works from errors in the details of genuine publications. Read Walters and Wilder in Scientific Reports.
Those percentages apply to the task, model versions and period studied. They do not measure AI errors in hospital maintenance, describe Arkmeds features or apply to current models. The relevant lesson for technical documentation is narrower: a plausible-looking reference does not establish that the source exists or is appropriate. When a suggestion mentions a standard, manual, procedure or tolerance, check the original source and its applicability before incorporating the information into the record.
The generative AI risk profile published by NIST in 2024 explicitly addresses false content presented with confidence. It recommends verifying sources and citations, documenting data provenance and evaluating capabilities empirically, while avoiding extrapolation from narrow examples. It is voluntary, cross-sector risk management guidance, rather than a certification of clinical engineering systems. Consult NIST AI 600-1.
To apply these principles to work orders, test the tool on cases where the original information is available and maintain the distinction between transcription and interpretation. Check for omitted outstanding issues, facts added without evidence, altered units and information linked to the wrong asset. Record which suggestions were accepted, corrected or rejected, along with review time. Together, these checks help establish whether the tool is useful in the actual workflow, rather than treating fluency or generation speed as evidence of quality.
Define who reviews the output and how discrepancies are handled
Assign responsibility for evaluating suggestions and define a process for handling inconsistencies. Reviewers must be able to reject a response, request more context or return a record for additional information.
Start with a small number of service records and compare the tool’s output with the original information. Record any problems found, such as an omitted outstanding issue, a change in the meaning of an observation or the inclusion of an undocumented fact.
Also assess whether the review fits the team’s routine. If a suggestion requires extensive correction, the use case, input data or presentation may need adjustment. Keep technical responsibility and equipment-related decisions within the applicable institutional process.
Measure usefulness against defined criteria
Choose indicators that fit the task. For documentation support, these might include records returned for additional information, suggestions corrected or rejected, and time spent reviewing. Define what each indicator means before comparing periods.
Changes in service volume or complexity can affect results. Describe the context and limits of any observed improvement, without turning one local experience into a promise for every operation.
Video on data and artificial intelligence
The Arkmeds video on the role of clinical engineering in the age of data and artificial intelligence, published on February 5, 2026, discusses data, management and decision support. It provides context for deciding which questions the team wants its data to answer. The video is in Brazilian Portuguese.
Frequently asked questions
Does AI perform equipment tests?
Measurements are performed by instruments. The AI functions discussed here support information processing and should be evaluated according to their role within the software.
Does more text in a work order mean higher quality?
Not necessarily. The record must be accurate, clear and sufficient to understand the service. Invented or repeated information does not improve documentation.
Can a suggestion be used without review?
Define a review process appropriate to the task and its risks. Content used in technical documentation needs to be checked before it is incorporated into the service conclusion.
Evaluate a use case in your operation
Talk to Arkmeds about the available features and how they fit your workflow. Describe the problem you want to solve, the data you have and your team’s review process.



