aehrc
/

cxrmate-rrg24

Feature Extraction

vision-encoder-decoder

Model card Files Files and versions Community

anicolson commited on Aug 26, 2024

Commit

c0566dd

·

verified ·

1 Parent(s): 09a7901

Update README.md

Files changed (1) hide show

README.md +3 -0

README.md CHANGED Viewed

@@ -101,6 +101,7 @@ There is no penalty in the reward for sampled reports that differ in length to t
 ## Citation:
 @inproceedings{nicolson-etal-2024-e,
     title = "e-Health {CSIRO} at {RRG}24: Entropy-Augmented Self-Critical Sequence Training for Radiology Report Generation",
     author = "Nicolson, Aaron  and
@@ -122,3 +123,5 @@ There is no penalty in the reward for sampled reports that differ in length to t
     pages = "99--104",
     abstract = "The core novelty of our approach lies in the addition of entropy regularisation to self-critical sequence training. This helps maintain a higher entropy in the token distribution, preventing overfitting to common phrases and ensuring a broader exploration of the vocabulary during training, which is essential for handling the diversity of the radiology reports in the RRG24 datasets. We apply this to a multimodal language model with RadGraph as the reward. Additionally, our model incorporates several other aspects. We use token type embeddings to differentiate between findings and impression section tokens, as well as image embeddings. To handle missing sections, we employ special tokens. We also utilise an attention mask with non-causal masking for the image embeddings and a causal mask for the report token embeddings.",
 }

 ## Citation:
+```
 @inproceedings{nicolson-etal-2024-e,
     title = "e-Health {CSIRO} at {RRG}24: Entropy-Augmented Self-Critical Sequence Training for Radiology Report Generation",
     author = "Nicolson, Aaron  and
     pages = "99--104",
     abstract = "The core novelty of our approach lies in the addition of entropy regularisation to self-critical sequence training. This helps maintain a higher entropy in the token distribution, preventing overfitting to common phrases and ensuring a broader exploration of the vocabulary during training, which is essential for handling the diversity of the radiology reports in the RRG24 datasets. We apply this to a multimodal language model with RadGraph as the reward. Additionally, our model incorporates several other aspects. We use token type embeddings to differentiate between findings and impression section tokens, as well as image embeddings. To handle missing sections, we employ special tokens. We also utilise an attention mask with non-causal masking for the image embeddings and a causal mask for the report token embeddings.",
 }
+```