Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Speech Recognition Software”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Use of speech recognition software: a vocal endurance test for the new millennium?

Speech recognition software for the personal or office computer is a relatively new area of technology. As the number of these products has increased so has use of this software. Some individuals will employ speech recognition systems due to difficulty with the conventional keyboard and mouse interface: others will use it for perceived efficiency or simply novelty. Regardless of the reason for use of this technology, the voice demands associated with extended or frequent use can be high, placing the user at risk for vocal difficulties. This paper reviews the case of an individual referred to our multidisciplinary voice care program for evaluation and treatment of vocal difficulties that began secondary to utilization of speech recognition software. We discuss medical and vocal histories, examination findings, treatment, and treatment outcomes.

Adult↗

Speech recognition software.

This article discusses the use of speech recognition software by means of reviewing two leading packages. Both programs require considerable training before they can be used effectively, but are then able to convert continuous speech into text with varying degrees of success.

CD-ROM↗

Comparative evaluation of three continuous speech recognition software packages in the generation of medical reports.

OBJECTIVE: To compare out-of-box performance of three commercially available continuous speech recognition software packages: IBM ViaVoice 98 with General Medicine Vocabulary; Dragon Systems NaturallySpeaking Medical Suite, version 3.0; and L&H Voice Xpress for Medicine, General Medicine Edition, version 1.2. DESIGN: Twelve physicians completed minimal training with each software package and then dictated a medical progress note and discharge summary drawn from actual records. MEASUREMENTS: Errors in recognition of medical vocabulary, medical abbreviations, and general English vocabulary were compared across packages using a rigorous, standardized approach to scoring. RESULTS: The IBM software was found to have the lowest mean error rate for vocabulary recognition (7.0 to 9.1 percent) followed by the L&H software (13.4 to 15.1 percent) and then Dragon software (14.1 to 15.2 percent). The IBM software was found to perform better than both the Dragon and the L&H software in the recognition of general English vocabulary and medical abbreviations. CONCLUSION: This study is one of a few attempts at a robust evaluation of the performance of continuous speech recognition software. Results of this study suggest that with minimal training, the IBM software outperforms the other products in the domain of general medicine; however, results may vary with domain. Additional training is likely to improve the out-of-box performance of all three products. Although the IBM software was found to have the lowest overall error rate, successive generations of speech recognition software are likely to surpass the accuracy rates found in this investigation.

Evaluation Studies as Topic↗

Combining speech recognition software with Digital Imaging and Communications in Medicine (DICOM) workstation software on a Microsoft Windows platform.

This presentation describes our experience in combining speech recognition software, clinical review software, and other software products on a single computer. Different processor speeds, random access memory (RAM), and computer costs were evaluated. We found that combining continuous speech recognition software with Digital Imaging and Communications in Medicine (DICOM) workstation software on the same platform is feasible and can lead to substantial savings of hardware cost. This combination optimizes use of limited workspace and can improve radiology workflow.

Humans↗

Speech recognition technology: an outlook for human-to-machine interaction.

Speech recognition, as an enabling technology in healthcare-systems computing, is a topic that has been discussed for quite some time, but is just now coming to fruition. Traditionally, speech-recognition software has been constrained by hardware, but improved processors and increased memory capacities are starting to remove some of these limitations. With these barriers removed, companies that create software for the healthcare setting have the opportunity to write more successful applications. Among the criticisms of speech-recognition applications are the high rates of error and steep training curves. However, even in the face of such negative perceptions, there remains significant opportunities for speech recognition to allow healthcare providers and, more specifically, physicians, to work more efficiently and ultimately spend more time with their patients and less time completing necessary documentation. This article will identify opportunities for inclusion of speech-recognition technology in the healthcare setting and examine major categories of speech-recognition software--continuous speech recognition, command and control, and text-to-speech. We will discuss the advantages and disadvantages of each area, the limitations of the software today, and how future trends might affect them.

Health Care Sector↗

Laboratory voice data entry system.

We have assembled a system using a personal computer workstation equipped with standard office software, an audio system, speech recognition software and an inexpensive radio-based wireless microphone that permits laboratory workers to enter or modify data while performing other work. Speech recognition permits users to enter data while their hands are holding equipment or they are otherwise unable to operate a keyboard. The wireless microphone allows unencumbered movement around the laboratory without a "tether" that might interfere with equipment or experimental procedures. To evaluate the potential of voice data entry in a laboratory environment, we developed a prototype relational database that records the disposal of radionuclides and/or hazardous chemicals. Current regulations in our laboratory require that each such item being discarded must be inventoried and documents must be prepared that summarize the contents of each container used for disposal. Using voice commands, the user enters items into the database as each is discarded. Subsequently, the program prepares the required documentation.

Computers↗

Speech recognition interface to a hospital information system using a self-designed visual basic program: initial experience.

Speech recognition (SR) in the radiology department setting is viewed as a method of decreasing overhead expenses by reducing or eliminating transcription services and improving care by reducing report turnaround times incurred by transcription backlogs. The purpose of this study was to show the ability to integrate off-the-shelf speech recognition software into a Hospital Information System in 3 types of military medical facilities using the Windows programming language Visual Basic 6.0 (Microsoft, Redmond, WA). Report turnaround times and costs were calculated for a medium-sized medical teaching facility, a medium-sized nonteaching facility, and a medical clinic. Results of speech recognition versus contract transcription services were assessed between July and December, 2000. In the teaching facility, 2042 reports were dictated on 2 computers equipped with the speech recognition program, saving a total of US dollars 3319 in transcription costs. Turnaround times were calculated for 4 first-year radiology residents in 4 imaging categories. Despite requiring 2 separate electronic signatures, we achieved an average reduction in turnaround time from 15.7 hours to 4.7 hours. In the nonteaching facility, 26600 reports were dictated with average turnaround time improving from 89 hours for transcription to 19 hours for speech recognition saving US dollars 45500 over the same 6 months. The medical clinic generated 5109 reports for a cost savings of US dollars 10650. Total cost to implement this speech recognition was approximately US dollars 3000 per workstation, mostly for hardware. It is possible to design and implement an affordable speech recognition system without a large-scale expensive commercial solution.

Cost-Benefit Analysis↗

Muscle tension dysphonia in patients who use computerized speech recognition systems.

The use of speech recognition systems as a replacement for other types of transcription systems is increasing rapidly, partly because many people are unable to use conventional keyboards as a result of upper-extremity repetitive strain injury (RSI). However, the frequent or continuous use of such systems can cause muscle tension dysphonia in some patients. The scientific literature suggests that there is an association between upper-extremity RSI and muscle tension dysphonia. We present a retrospective case series of five patients with workplace upper-extremity RSI who developed muscle tension dysphonia soon after they began using discrete computerized speech recognition software. The diagnosis of dysphonia was based on laryngovideostroboscopy, acoustic analyses, and voice load testing. All patients had normal voice when using everyday speech, but speaking into the computer resulted in the rapid onset of aperiodicity, strain, and a decrease in fundamental frequency. In three of the five patients, laryngovideostroboscopy showed posterior glottic overapproximation, but no other abnormalities. Treatment was centered on voice therapy and avoidance of long periods of using computerized speech recognition systems. The condition of three of the five patients improved with therapy. We conclude that computer speech recognition programs can lead to the onset of muscle tension dysphonia in some patients. These patients can be successfully treated with voice therapy.

Adult↗

Automatic concept extraction from spoken medical reports.

OBJECTIVE: The objective of this project is to investigate methods whereby a combination of speech recognition and automated indexing methods substitute for current transcription and indexing practices. METHODS: We based our study on existing speech recognition software programs and on NOMINDEX, a tool that extracts MeSH concepts from medical text in natural language and that is mainly based on a French medical lexicon and on the UMLS. For each document, the process consists of three steps: (1) dictation and digital audio recording, (2) speech recognition, (3) automatic indexing. The evaluation consisted of a comparison between the set of concepts extracted by NOMINDEX after the speech recognition phase and the set of keywords manually extracted from the initial document. The method was evaluated on a set of 28 patient discharge summaries extracted from the MENELAS corpus in French, corresponding to in-patients admitted for coronarography. RESULTS: The overall precision was 73% and the overall recall was 90%. Indexing errors were mainly due to word sense ambiguity and abbreviations. A specific issue was the fact that the standard French translation of MeSH terms lacks diacritics. A preliminary evaluation of speech recognition tools showed that the rate of accurate recognition was higher than 98%. Only 3% of the indexing errors were generated by inadequate speech recognition. DISCUSSION: We discuss several areas to focus on to improve this prototype. However, the very low rate of indexing errors due to speech recognition errors highlights the potential benefits of combining speech recognition techniques and automatic indexing.

Abstracting and Indexing↗

Experimental analysis of human vocal behavior: applications of speech-recognition technology.

Recent developments in speech recognition make it feasible to apply the technology to study vocal behavior. The present study illustrates the use of this technology to establish functional stimulus classes. Eight students were taught to say nonsense words in the presence of arbitrarily assigned sets of symbols consistent with three three-member experimenter-defined stimulus classes. Computer-controlled speech-recognition software was used to record, analyze, and differentially reinforce vocal responses. When the stimulus classes were established, students were taught to say a new nonsense word in the presence of one member of each stimulus class. Transfer of function was tested subsequently to determine if the novel stimulus names transferred to the remaining stimulus class members. Most subjects required two iterations of the training and testing procedures before transfer occurred. The data illustrate the usefulness of recording vocal behavior during stimulus control procedures and demonstrate the use of speech-recognition technology. The paper also describes the current state of speech-recognition technology and suggests several other areas of research that might benefit from using vocal behavior as its primary datum.

Adolescent↗

Automated speech recognition for time recording in out-of-hospital emergency medicine-an experimental approach.

Precise documentation of medical treatment in emergency medical missions and for resuscitation is essential from a medical, legal and quality assurance point of view [Anästhesiologie und Intensivmedizin, 41 (2000) 737]. All conventional methods of time recording are either too inaccurate or elaborate for routine application. Automated speech recognition may offer a solution. A special erase programme for the documentation of all time events was developed. Standard speech recognition software (IBM ViaVoice 7.0) was adapted and installed on two different computer systems. One was a stationary PC (500MHz Pentium III, 128MB RAM, Soundblaster PCI 128 Soundcard, Win NT 4.0), the other was a mobile pen-PC that had already proven its value during emergency missions [Der Notarzt 16, p. 177] (Fujitsu Stylistic 2300, 230Mhz MMX Processor, 160MB RAM, embedded soundcard ESS 1879 chipset, Win98 2nd ed.). On both computers two different microphones were tested. One was a standard headset that came with the recognition software, the other was a small microphone (Lavalier-Kondensatormikrofon EM 116 from Vivanco), that could be attached to the operators collar. Seven women and 15 men spoke a text with 29 phrases to be recognised. Two emergency physicians tested the system in a simulated emergency setting using the collar microphone and the pen-PC with an analogue wireless connection. Overall recognition was best for the PC with a headset (89%) followed by the pen-PC with a headset (85%), the PC with a microphone (84%) and the pen-PC with a microphone (80%). Nevertheless, the difference was not statistically significant. Recognition became significantly worse (89.5% versus 82.3%, P<0.0001 ) when numbers had to be recognised. The gender of speaker and the number of words in a sentence had no influence. Average recognition in the simulated emergency setting was 75%. At no time did false recognition appear. Time recording with automated speech recognition seems to be possible in emergency medical missions. Although results show an average recognition of only 75%, it is possible that missing elements may be reconstructed more precisely. Future technology should integrate a secure wireless connection between microphone and mobile computer. The system could then prove its value for real out-of-hospital emergencies.

Automation↗

[Experiences with a current speech recognition system in creating cardiology reports].

Development of speech recognition software is at a stage where you can use it effectively for creating cardiological reports or at least parts of it. We were very successful in using the Dragon Naturally Speaking system for our reports and we don't need a secretary for our writings any longer. Very important for effective work is a fast and good PC hardware, especially a good sound system and a sufficient amount of internal memory, also important is patience of the user, because a longer training phase is required. I recommend for own experiences to use the standard version which is cheap and includes all necessary features. Following some fundamental rules anyone can be successful with speech recognition. Development of speech recognition is increasing rapidly so that everyone of us will get in contact with it sooner or later, it would be better to be prepared for it now.

Cardiology↗

[Edeka does all--machine speech recognition in social medicine expert testimony].

UNLABELLED: Automatic speech recognition systems are already being used in spheres employing a restricted vocabulary. OBJECTIVE: Our aim was to investigate whether low-cost speech recognition software for PC is capable of being usefully employed in the sphere of sociomedicine. MATERIALS AND METHODS: To this end 34 representative pages of text (a total of 11,000 words) taken from expertises on cases of suspected medical malpractice (many different subspecialties) were dictated using IBM's "Voice Type Simply Speaking" software. Having completed a page, the resulting error rate was recorded, and the text was corrected before we proceeded with the dictation. Finally, 3 pages of text were re-dictated and the resulting error rate determined. RESULTS: The error rate in the previously unknown text ranged between 10 and 23 per cent (mean 15.9%) without any significant reduction during the training phase, while that in the re-dictated text was drastically reduced to less than 3 per cent. It became evident that once a word was corrected the system hardly ever repeated that particular mistake. CONCLUSION: The system's poor performance on unknown text and the missing reduction in the error rate during the training phase are obviously not due to any incompetence of the system but to the huge amount of technical jargon in the scope of medical writing. To attain an acceptable performance we suggest to either extend the training phase, or, preferably, to confine the application to a single medical subspecialty. Its overwhelming learning ability makes the system a serious candidate typist in the sphere of sociomedicine.

Expert Testimony↗

Importance and effects of altered workplace ergonomics in modern radiology suites.

The transition from a film-based to a filmless soft-copy picture archiving and communication system (PACS)-based environment has resulted in improved work flow as well as increased productivity, diagnostic accuracy, and job satisfaction. Adapting to this filmless environment in an efficient manner requires seamless integration of various components such as PACS workstations, the Internet and hospital intranet, speech recognition software, paperless electronic hospital medical records, e-mail, office software, and telecommunications. However, the importance of optimizing workplace ergonomics has received little attention. Factors such as the position of the work chair, workstation table, keyboard, mouse, and monitors, along with monitor refresh rates and ambient room lighting, have become secondary considerations. Paying close attention to the basics of workplace ergonomics can go a long way in increasing productivity and reducing fatigue, thus allowing full realization of the potential benefits of a PACS. Optimization of workplace ergonomics should be considered in the basic design of any modern radiology suite.

Asthenopia↗

[Creating language model of the forensic medicine domain for developing a autopsy recording system by automatic speech recognition].

For the purpose of practical use of speech recognition technology for recording of forensic autopsy, a language model of the speech recording system, specialized for the forensic autopsy, was developed. The language model for the forensic autopsy by applying 3-gram model was created, and an acoustic model for Japanese speech recognition by Hidden Markov Model in addition to the above were utilized to customize the speech recognition engine for forensic autopsy. A forensic vocabulary set of over 10,000 words was compiled and some 300,000 sentence patterns were made to create the forensic language model, then properly mixing with a general language model to attain high exactitude. When tried by dictating autopsy findings, this speech recognition system was proved to be about 95% of recognition rate that seems to have reached to the practical usability in view of speech recognition software, though there remains rooms for improving its hardware and application-layer software.

Autopsy↗

Review: occupational risks for voice problems.

The purpose of this paper is to provide a cohesive review of the literature regarding the functional consequences of voice problems and occupational risk factors for them. The salient points are as follows. According to conservative estimates, approximately 28,000,000 workers in the US experience daily voice problems. Many people who experience voice problems perceive them to have a negative impact on their work and their quality of life. Estimates based on empirical data suggest that, considering only lost work days and treatment expenses, the societal cost of voice problems in teachers alone may be of the order of about $2.5 billion annually in the US. In fact, across several countries, "teacher" consistently emerges as the common occupation most likely to seek otorhinolaryngological (ORL) evaluation for a voice problem. Other occupational categories likely to seek ORL examination for a voice problem are singer, counselor/social worker, lawyer, and clergy. Finally, US Census data indicate that keyboard operators may be at special risk for voice problems because of a near-epidemic growth of repetitive strain injury (RSI), which can adversely affect the voice especially when speech recognition software is implemented. This paper discusses frequency data, quality of life data, and treatment considerations for these voice-related occupational issues.

Adult↗