Conversation
dustin-sale
left a comment
There was a problem hiding this comment.
Review by @dustin-sale via Codex.
Requested changes
src/oci-speech-mcp-server/oracle/oci_speech_mcp_server/utils/responses.py:129— [P1] Write local artifacts privately and atomically.src/oci-speech-mcp-server/oracle/oci_speech_mcp_server/tools/transcription.py:468— [P1] Preserve or clean uploaded media when job creation fails.
Additional review notes
- [P2] Surface created notification resources after partial failure.
- [P2] Validate model-specific transcription options.
- [P2] Preserve actionable local I/O errors.
Validation
make test project=oci-speech-mcp-serverpassed: 45 tests and 97.21% coverage.make lintand all reported GitHub checks passed. Live OCI operations were not rerun.
See the inline comments in this review for evidence, impact, and suggested remediation.
- write local outputs privately and atomically - clean uploaded media when transcription job creation fails - return structured state after partial notification or cleanup failures - validate Oracle and Whisper model-specific options - preserve actionable, sanitized local I/O diagnostics - add regression tests and update documentation
dustin-sale
left a comment
There was a problem hiding this comment.
Review by @dustin-sale via Codex.
Requested changes
src/oci-speech-mcp-server/oracle/oci_speech_mcp_server/tools/transcription.py:37— [P1] Recognize supported Whisper model identifiers.src/oci-speech-mcp-server/oracle/oci_speech_mcp_server/tools/transcription.py:475— [P1] Preserve recovery state after side effects begin.
Additional review notes
src/oci-speech-mcp-server/oracle/oci_speech_mcp_server/__init__.py:8— [P2] Derive the package version from distribution metadata.
Validation
make lintpassed.make test project=oci-speech-mcp-serverpassed with 47 tests and 96.54% coverage; all current GitHub checks also pass.
See the inline comments in this review for evidence, impact, and suggested remediation.
| known_model = model_type.strip().upper() | ||
| if known_model == "ORACLE" and whisper_prompt: | ||
| raise ValueError("whisper_prompt is supported only when model_type=WHISPER.") | ||
| if known_model == "WHISPER" and not punctuation_enabled: | ||
| raise ValueError("punctuation_enabled must be true when model_type=WHISPER.") |
There was a problem hiding this comment.
Comment from @dustin-sale via Codex.
[P1] Recognize supported Whisper model identifiers
Evidence: This validation only recognizes the exact value WHISPER, but OCI SDK 2.182.1 lists WHISPER_MEDIUM and WHISPER_LARGE_V2; current OCI documentation also lists WHISPER_LARGE_V3_TURBO. A local probe showed that WHISPER_MEDIUM passes this function with punctuation_enabled=False and is then constructed with the Oracle default language_code="en-US". OCI documents locale-agnostic Whisper codes such as en and auto, and requires punctuation for Whisper. See https://docs.oracle.com/en-us/iaas/tools/python/latest/api/ai_speech/models/oci.ai_speech.models.TranscriptionModelDetails.html and https://docs.oracle.com/iaas/Content/speech/using/create-trans-job.htm.
Impact: Requests using the supported Whisper identifiers bypass the intended model-specific validation, so ordinary defaults or disabled punctuation can create asynchronously failed jobs instead of returning an immediate actionable error.
Requested change: Recognize the actual Whisper-family identifiers, enforce their language-code and punctuation rules while retaining intentional future-model extensibility, and replace the WHISPER tests with cases using at least WHISPER_MEDIUM and WHISPER_LARGE_V2.
| normalization=normalization(punctuation_enabled, profanity_mode), | ||
| additional_transcription_formats=["SRT"] if include_srt else [], | ||
| ) | ||
| with source.open("rb") as media: |
There was a problem hiding this comment.
Comment from @dustin-sale via Codex.
[P1] Preserve recovery state after side effects begin
Evidence: The generated Object Storage identity exists before this upload, but only a later create_transcription_job failure has special partial-state handling. An ambiguous put_object exception, or any polling, task lookup, output download, or local-write failure after job creation, reaches the outer raise_safe handler without returning the generated object name, upload location, or created job ID.
Impact: A PUT accepted before a lost response can leave sensitive media in Object Storage without a cleanup identity. A failure after job creation similarly prevents reliable resume, inspection, and cleanup, especially because display_name is not guaranteed to be unique.
Requested change: Track the workflow stage and safe partial state from before upload. On upload failure, attempt an idempotent delete or return the object identity and cleanup status. After job creation, return the upload and job identifiers with sanitized failure metadata. Add regressions for an ambiguous upload failure and a post-create download failure.
| """ | ||
|
|
||
| __project__ = "oracle.oci-speech-mcp-server" | ||
| __version__ = "1.0.0" |
There was a problem hiding this comment.
Comment from @dustin-sale via Codex.
[P2] Derive the package version from distribution metadata
Evidence: __version__ = "1.0.0" duplicates the version in pyproject.toml. BEST_PRACTICES.md requires importlib.metadata.version(__project__), and the other Python servers in this checkout follow that pattern. The additional OCI user agent derives from this literal, while its test separately hardcodes oci-speech-mcp/1.0.0.
Impact: A future package-version bump can silently continue reporting stale OCI SDK telemetry, and the current test would still pass.
Requested change: Derive __version__ from installed distribution metadata without a fallback, and have the user-agent test derive its expected version from the same package metadata rather than another literal.
There was a problem hiding this comment.
@Prabhutva we made a recent update to simplify derivation for package versions. Example:
from importlib.metadata import version as distribution_version
__project__ = "oracle.oci-speech-mcp-server"
__version__ = distribution_version(__project__)
Description
Adds the initial release of the OCI Speech MCP Server (
v1.0.0), a locally runstdioMCP server that enables agents and MCP clients to use OCI Speech through the OCI Python SDK.Transcription Jobs
Adds the following tools:
create_transcription_jobget_transcription_joblist_transcription_jobsupdate_transcription_jobdelete_transcription_jobcancel_transcription_jobchange_transcription_job_compartmentlist_transcription_tasksget_transcription_taskcancel_transcription_taskdownload_transcription_resultstranscribe_local_fileThe
transcribe_local_filetool provides an end-to-end local transcription workflow. It validates and uploads a local media file to OCI Object Storage, creates and optionally monitors the transcription job, and downloads successful JSON and SRT outputs to a restricted local directory.Transcription features include diarization, configurable speaker counts, punctuation, profanity filtering, multiple transcription models and domains, Whisper prompting, and SRT output.
Resources and prompts provide guidance for:
Customizations
Adds the following tools:
create_customizationget_customizationlist_customizationsupdate_customizationdelete_customizationchange_customization_compartmentCustomization inputs support inline entities, pronunciations, reference examples, Object Storage datasets, and reusable entity customizations.
The accompanying resource and prompt explain how to select, structure, create, train, update, and reuse OCI Speech customizations.
Text to Speech
Adds the following tools:
list_voicessynthesize_speechSpeech synthesis supports plain text and SSML, configurable voices, audio formats, sample rates, and safe local output handling.
The TTS and SSML resources include examples and guidance for:
Extras
Adds the following tool:
setup_transcription_notificationsThis tool creates or reuses an OCI Notifications topic and creates OCI Events rules for transcription job completion and failure events. Subscription creation remains an explicit user action so confirmation endpoints are not configured without consent.
Additional resources and prompts cover:
Realtime transcription guidance keeps the persistent WebSocket connection in the user's application, where audio capture and playback occur, instead of holding a long-running connection inside the
stdioMCP server.Motivation and context
This server makes OCI Speech workflows directly accessible to MCP-compatible agents while continuing to use the user's existing OCI credentials and authorization policies.
It reduces the setup required for workflows involving Object Storage, transcription result retrieval, diarization, customizations, SSML generation, speech synthesis, and transcription job notifications.
The implementation separates tools, prompts, and resources by capability so future OCI Speech modules can be added without expanding the server entry point.
The server also provides:
oracle-mcp-commonoci-speech-mcp/1.0.0OCI SDK user agentDependencies and prerequisites
Runtime dependencies include:
fastmcp==3.4.5oci==2.182.1oracle-mcp-common>=0.1.0,<0.2.0pydantic>=2.13.4,<3Users require a configured OCI authentication profile and the appropriate OCI Speech, Object Storage, Events, Notifications, and IAM permissions for the operations they intend to use.
Type of change
How Has This Been Tested?
The server was validated with:
make test project=oci-speech-mcp-server make lintAutomated validation results:
oci-speech-mcp/1.0.0user agentLive
stdioMCP validation confirmed discovery of all 21 tools, 11 resources, and 9 prompts. Every resource was read and every prompt was rendered.Transcription Jobs
Customizations
Text to Speech
Extras
To reproduce the live tests:
stdio.Test Configuration:
us-phoenix-1Checklist: